Multiturn Conversation and Prompts for Structured Analysis of CAR-T Cell Therapy R&D Documents

R&D document data sources in CAR-T cell therapy include clinical trial reports, patent literature, academic papers, internal experimental records, and

Data Characteristics

R&D document data sources in CAR-T cell therapy include clinical trial reports, patent literature, academic papers, internal experimental records, and regulatory submission materials. These documents primarily exist as PDFs, Word files, or in structured databases. Data updates frequently, especially clinical trial progress and new drug application information. Document structures are complex, containing extensive specialized terminology, abbreviations, gene sequences, protein structures, cell line information, clinical indicators, and dosage units. For example, cell line names may include special characters and number combinations. Clinical trial data involves patient cohorts, adverse event rates, efficacy evaluation metrics (such as complete response (CR) rate, partial response (PR) rate), and concentration units (nM, µg/mL) and time units (days, weeks, months).

Constraints Imposed by These Characteristics on Multiturn Conversation and Prompts

The complexity and specialized nature of CAR-T cell therapy documents impose specific requirements on the design of multiturn conversations and prompts. The use of specialized terminology and abbreviations requires prompts to guide the model in precise contextual understanding and disambiguation. For instance, an abbreviation might represent different biological entities or clinical indicators in different contexts. High data update frequency means the conversation system must index and cite the latest information promptly, avoiding outdated content. Specific data modalities within documents, such as gene sequences and protein structures, require the model to identify and process this non-textual information, and accurately present it in conversations. Additionally, numerical values and units in clinical trial data require the model to perform accurate numerical comparisons and unit conversions after structured analysis, preventing misinterpretation due to unit confusion.

Configuration Settings

Configuration ItemSuggested ValueRationale
maxContext6 turnsBalances contextual coherence with computational resource consumption, covering typical analysis workflows.
Chunk size800 charactersAdapts to paragraph lengths in complex biomedical texts, ensuring semantic completeness.
Recall count10 itemsIncreases coverage of relevant information, raising the probability of finding key details in specialized documents.
Similarity threshold0.75Precisely matches highly specialized terminology and concepts, reducing interference from irrelevant information.
Rerank result count5 itemsFurther focuses on the most relevant document segments, optimizing model processing efficiency.
temperature0.3Ensures the rigor and accuracy of model output, reducing the risk of hallucination.

Common Pitfalls

  • The model frequently encounters specialized terms or abbreviations in conversations and fails to correctly understand their meaning in the current context, leading to off-topic or factually incorrect answers. This occurs because prompts do not explicitly require the model to perform contextual disambiguation of specialized terms.
  • The model cites outdated clinical trial data or drug information in multiturn conversations, resulting in inaccurate analysis conclusions. This happens when the knowledge base indexing update mechanism fails to synchronize with the latest R&D progress in a timely manner.
  • In questions involving numerical comparisons or unit conversions, the model provides results that do not match actual data or have incorrect units. This is due to the failure to accurately extract and standardize numerical values and their unit information during document structured analysis.

Validation of Configuration

  • Create a test question set with different contexts for core specialized terms and abbreviations. Check if the model consistently understands and references them correctly in multiturn conversations.
  • Regularly query the system about the latest clinical trial progress or drug submission status. Verify if the model can cite the most current knowledge base content and cross-reference the publication date with the model's citation time.
  • Design query scenarios involving numerical comparisons and unit conversions, such as "compare the difference in CR rates between two CAR-T products" or "convert a concentration from nM to µg/mL." Verify the accuracy of the numerical values and units provided by the model.

Note: The values provided are common starting points and should be measured against specific samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.