Data Characteristics for This Category
CAR-T cell therapy registration dossiers involve multiple data sources. Clinical trial data primarily originates from hospital Electronic Medical Record (EMR) systems and clinical research databases. During the trial phase, updates typically occur weekly. For the submission phase, a final locked version is used. Manufacturing process data comes from pharmaceutical companies' internal Manufacturing Execution Systems (MES) and Quality Management Systems (QMS). Data updates are relatively stable, occurring mainly during process optimization or changes. Non-clinical research data often consists of experimental reports in PDF or Word formats. Literature data is sourced from specialized databases like PubMed and Web of Science, with a higher frequency of updates. These dossier documents have complex structures, containing numerous tables, figures, and cross-references. Fields include specific terminology such as dosage units (e.g., cells/kg), biomarkers (e.g., CD19+ cell count), and adverse events (e.g., CRS grade).
Constraints Imposed by These Characteristics on Multi-Turn Conversations and Prompts
The data characteristics of CAR-T cell therapy dossiers impose specific requirements on the configuration of multi-turn conversations and prompts. First, complex document structures and numerous figures make it difficult for traditional text chunking to effectively capture context, necessitating more refined text preprocessing and embedding strategies. Second, unique biomedical terminology and units require the model to have stronger domain understanding. Prompt design must clearly define terminology to avoid ambiguity. The real-time nature of clinical data requires the ability to dynamically update or reference the latest data versions within multi-turn conversations. Heterogeneous multi-source data means that during a conversation, it may be necessary to simultaneously retrieve and integrate fragments from different systems. Prompts must guide the model to perform cross-document information aggregation. Furthermore, strict scrutiny of safety and efficacy mandates that conversations precisely cite original data and literature, requiring extremely high recall accuracy.
Configuration Recommendations
| Configuration Item | Recommended Value | Rationale for Recommendation |
|---|---|---|
Chunk size (Chunk Length) | 500 characters (500 characters) | Balances semantic completeness of long texts with retrieval efficiency, preventing loss of critical information due to splitting, especially for clinical trial reports. |
Recall count (Recall Count) | 8 entries (8 items) | Considering the complexity and multi-dimensional information associations in CAR-T dossiers, increasing the recall count appropriately covers more relevant context. |
Similarity threshold (Similarity Threshold) | 0.82 | A high threshold ensures that recalled text snippets are highly relevant to the query, reducing interference from irrelevant information, which is crucial for the rigor of registration submissions. |
maxContext | 6000 Token | Allows for longer conversation history and recalled content, supporting complex multi-turn follow-up questions and information integration, addressing cross-chapter references. |
Rerank result count (Reranked Return Count) | 5 entries (5 items) | Reranks recall results to prioritize the most critical evidence, improving the precision of conversational responses. |
LLM Model | ERNIE-4.0 | Selects a large model with strong Chinese comprehension and logical reasoning capabilities to more accurately process specialized biomedical terminology and complex sentence structures. |
Three Common Pitfalls
- Critical data missing or incorrect citations appear in conversations because of improper knowledge base chunking strategies, leading to truncation of key numbers or units.
- The model's interpretation of biomarkers or adverse events is inconsistent across multi-turn conversations because prompts do not explicitly define the scope of contextual interpretation for domain-specific terms.
- The model cannot provide specific values when users ask follow-up questions about particular clinical trial results because tabular data in original documents was not structured, making it irretrievable.
How to Verify Configuration
- Conduct multi-turn conversation tests for core registration submission questions. Check the accuracy of data sources cited by the model and compare them with original documents to verify the reasonableness of the
Similarity threshold(Similarity Threshold). - Randomly select complex tabular data from clinical trial reports. Ask questions about relevant values in the conversation and check the precision and completeness of the model's returned results to evaluate the effectiveness of
Chunk size(Chunk Length). - Simulate the questioning path of a submission reviewer. Observe whether the model's explanation of specific CAR-T adverse events (e.g.,
CRS,ICANS) remains consistent across multi-turn conversations and check the capacity ofmaxContext.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.