Data Characteristics
Hemato-oncology regulatory submission documents typically include clinical trial protocols, investigator brochures, clinical study reports (CSRs), chemistry, manufacturing, and control (CMC) data, and non-clinical study reports. Document types vary, including PDF, Word, and Excel. Data update frequency is relatively low, primarily occurring with the release of different interim and final reports for clinical trials. Document structures are complex, containing numerous tables, figures, and cross-references. Fields and units are highly specialized, such as dosage units like mg/kg, efficacy evaluation criteria like CR (Complete Remission) and PR (Partial Remission), and specific gene mutation sites like FLT3-ITD and IDH1/2. Clinical trial data often appear in structured table format, but with extensive unstructured text descriptions.
Constraints Imposed by These Characteristics on Multi-turn Conversations and Prompts
The complexity of hemato-oncology regulatory submission documents challenges the effectiveness of multi-turn conversations and prompts. The large volume of specialized terminology and abbreviations in documents requires the model to possess deep domain knowledge to avoid semantic misunderstandings. Cross-references and tabular data make it difficult for a single context window to capture complete information, necessitating longer context management and cross-document retrieval capabilities. The low data update frequency means that knowledge base construction must focus on the consistency and accuracy of historical data, ensuring dialogue results are based on the latest approved submission versions. Accurate identification and parsing of specialized fields and units directly impact the reliability of the question-answering system. For example, any deviation in dosage calculation or efficacy determination could lead to serious consequences.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
maxContext | 8000 tokens | Accommodates lengthy clinical study reports and detailed pharmaceutical descriptions in hemato-oncology submissions, ensuring context completeness. |
Chunk size (Segment Length) | 500 characters (characters) | Balances semantic integrity and retrieval efficiency. Avoids introducing irrelevant information due to overly long segments while retaining sufficient context. |
Recall count (Recall Count) | Top 8 entries (top 8) | Hemato-oncology data is highly interconnected. Increasing the recall count helps cover more potentially relevant information, improving answer accuracy. |
Similarity threshold (Similarity Threshold) | 0.75 | Ensures recalled results are highly relevant to the user's query, filtering out document snippets that are medically similar in terminology but irrelevant in content. |
Rerank result count (Rerank Return Count) | 3 entries (3 items) | After a high recall count, reranking selects the few most relevant items, reducing the model's processing burden and improving response speed. |
Historical Conversation Turns | 5 Turns (5 turns) | Supports multi-turn follow-up questions and detail confirmation, especially in complex scenarios such as dosage adjustments and adverse event analysis. |
Common Pitfalls
- Slow conversation response, exceeding
10 seconds(10 seconds): Often caused by an excessively largemaxContextor too manyRecall count(Recall Count), leading to excessive context processing and retrieval pressure on the model. - Answers missing critical data or inaccurate units: Occurs when
Chunk size(Segment Length) is too short, truncating important numerical information or context, or when specialized fields are not adequately identified and extracted in the knowledge base. - Model unable to associate previous information during user follow-up questions: Caused by insufficient
Historical Conversation Turnssettings, leading the model to forget background information from previous turns.
How to Verify Configuration
- Conduct multi-turn dialogue tests with complex questions from typical submission documents. Observe whether answers accurately include specialized terminology, dosages, and units.
- Randomly select documents containing tabular data from the knowledge base. Ask queries involving table content and verify if the model correctly parses and references the data.
- Simulate a user's continuous questioning during a specific submission stage. Check if the model maintains coherent understanding of the topic across different turns and can infer based on previous information.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.