Data Characteristics in This Category
Hematologic oncology data originates from diverse sources, including clinical trial reports, gene sequencing results, pathology diagnostic reports, multi-center study data, and drug inserts. Data updates are frequent, especially with new drug approvals, clinical guideline revisions, and gene mutation discoveries, typically occurring quarterly or semi-annually. Document structures vary, encompassing unstructured clinical notes, semi-structured laboratory reports (e.g., CBC complete blood count, FISH fluorescence in situ hybridization results), and structured drug target information. Fields and units are highly specialized; for example, chromosomal translocation t(9;22), fusion gene BCR-ABL1, white blood cell count (WBC) with units of 10^9/L, and minimal residual disease (MRD) results expressed as percentages. Drug dosages are often given in mg/kg or mg/m^2, with multiple administration protocols.
Constraints Imposed by These Characteristics on Model Integration and Configuration
The high specialization and diversity of hematologic oncology data require robust text understanding and multimodal information processing capabilities during model integration. Frequent data updates necessitate regular incremental or full synchronization of the knowledge base to ensure the timeliness and accuracy of model outputs. The coexistence of unstructured and semi-structured data challenges the robustness of document parsers, requiring finely tuned preprocessing pipelines to extract key entities and numerical values. The accuracy of identifying specific fields (e.g., gene loci, drug names, clinical symptoms) directly impacts consultation quality, demanding targeted optimization of Named Entity Recognition (NER) models or enhancement with domain-specific dictionaries. Unit complexity requires the model to accurately convert and interpret units in its responses, avoiding confusion. Furthermore, the dynamic nature of disease progression and treatment plans dictates that knowledge retrieval strategies must balance breadth and depth to support complex clinical decision-making.
How to Set Configurations
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
maxContext | 32000 tokens | Hematologic oncology documents are often long, containing detailed clinical data and genetic information, requiring a larger context window. |
Chunk size (Segment Length) | 800–1200 characters (characters) | Ensures each segment contains sufficient semantic information while avoiding excessive length that could dilute key information, suitable for gene reports and pathological analyses. |
Recall count (Recall Count) | Top 8 entries (top 8) | Considering disease complexity and diverse treatment plans, more relevant document snippets are needed to assist model judgment. |
Similarity threshold (Similarity Threshold) | 0.78 | Increases recall precision and reduces interference from irrelevant information, especially when distinguishing similar symptoms or drugs. |
Rerank result count (Reranked Return Count) | Top 3 entries (top 3) | Further optimizes relevance using a reranking model, focusing on the most core diagnostic or treatment evidence. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds (seconds) | Parsing time can be long when processing large clinical research reports or multiple gene sequencing files. |
Three Common Pitfalls
- Symptom: Model responses contain incorrect drug dosages or gene locus information. Reason: The original document parsing failed to correctly identify numerical values with units or specific gene naming conventions, leading to inaccurate knowledge base indexing.
- Symptom: When users inquire about the latest treatment progress, the model's response is outdated. Reason: The knowledge base synchronization strategy is not executed regularly or the update frequency is insufficient, failing to incorporate new clinical trial data or guideline revisions in a timely manner.
- Symptom:
503error messageNo available channel for model yi-vl-puls under current group default. Reason: The modelyi-vl-pulsspecified in the model configuration is not correctly deployed in the backendOneAPIor other model service platforms, or the corresponding access channel is not enabled.
How to Confirm Correct Configuration
- Conduct multi-turn dialogue tests for typical hematologic oncology cases (e.g.,
CMLchronic myeloid leukemia,AMLacute myeloid leukemia). Check the accuracy of the model's responses regarding disease diagnosis, treatment plans, and prognosis, paying particular attention to the expression of gene mutations and drug dosages. - Upload a recent clinical trial report and a drug insert. Check if the document parser can correctly extract key fields such as
BLASTprimitive cell ratio,fusion genetype, anddrug dosage, and verify if they are successfully indexed in the knowledge base. - Simulate user inquiries involving specific
chromosomal translocation t(15;17)orFLT3-ITDmutations. Check if the model can recall relevant treatment guidelines or targeted drug information, and evaluate the relevance and completeness of the recalled documents.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.