Data Characteristics
Hematologic oncology data comes from various sources. These include clinical trial reports, drug inserts, academic papers, genetic testing reports, and patient case files. Document update frequency is high, especially for new drug development and clinical research. New data can appear weekly or even daily.
Document structures vary. Drug inserts and clinical guidelines typically have standardized chapters and paragraphs. Academic papers follow an abstract, introduction, methods, results, and discussion structure. Genetic testing reports contain numerous biomarker names, gene mutation sites, and expression levels, often presented in tables.
Fields and units include drug dosage (e.g., mg/kg), treatment duration (e.g., days, weeks), gene loci (e.g., Exon 12), mutation frequency (e.g., %), and clinical indicators (e.g., white blood cell count 10^9/L).
Constraints on Knowledge Base Retrieval and Recall
The high update frequency of hematologic oncology data requires efficient incremental updates and version management in the knowledge base. This ensures timely retrieval results.
Standardized document structures improve text segmentation accuracy and reduce semantic fragmentation. However, diverse document types require tailored parsing strategies.
Genetic testing reports contain many specialized terms and numerical data. This challenges the recall algorithm's precise matching capabilities. The system must recognize and associate medical terms expressed in different ways.
Diverse units for drug dosages and clinical indicators mean simple keyword matching is insufficient. Smarter retrieval needs to combine numerical ranges and unit conversions. This avoids recall failures due to unit differences.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size | 500–800 characters | Adapts to paragraph structures in clinical guidelines and academic papers, maintaining semantic completeness. |
Recall count | 8–12 entries | Covers potentially relevant information while balancing subsequent re-ranking efficiency. |
Similarity threshold | 0.75–0.85 | Balances recall precision and generalization ability, reducing interference from irrelevant results. |
Rerank result count | 3–5 entries | Focuses on the most relevant key information, reducing the model's processing burden. |
Embedding Model Version | text-embedding-ada-002 or newer | Ensures accuracy of semantic understanding and quality of vectorization. |
Update Frequency | Once daily | Addresses rapid updates in new drug development and clinical trial data. |
Common Mistakes
- Retrieval results do not include the latest clinical research data. This occurs when the knowledge base synchronization mechanism is not triggered in time or the data source update frequency is lower than required.
- Searching for specific gene mutation sites returns empty or irrelevant results. This happens when multiple expressions for gene names or locus numbers are not uniformly processed, leading to recall matching failure.
- Dynamically selecting a knowledge base in a workflow fails to pass the correct knowledge base ID. This prevents locating the target knowledge base for retrieval, and the system reports the knowledge base does not exist.
Verification Steps
- Import recent new drug clinical trial data into the knowledge base. Use relevant queries to verify if the latest information is accurately recalled and check the completeness of the recall results.
- Construct gene mutation queries with different expressions. For example, use both full gene names and abbreviations, or different locus number formats. Check if relevant reports are consistently recalled.
- Configure dynamic knowledge base selection in the workflow. Use system logs or debugging tools to check if the passed knowledge base ID matches the expectation. Verify that the retrieval operation executes successfully.
The values provided are common starting points and should be measured against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.