Data Characteristics
R&D documents in hematologic oncology are diverse and originate from multiple sources. Data primarily comes from clinical trial reports, pathology analysis reports, gene sequencing data, drug specifications, and related scientific literature. These documents update frequently; clinical trial data and gene sequencing results, in particular, may see new batches released weekly or even daily. Document structures vary: clinical reports often contain structured or semi-structured fields such as patient demographics, diagnostic results, treatment plans, and efficacy evaluations, while scientific literature is predominantly unstructured text. Common fields and units include ECOG scores, CR/PR/SD/PD efficacy evaluation criteria, Copy Number Variation (CNV), gene mutation frequency (%), and concentration units like μg/mL and nM.
Constraints Imposed by Data Characteristics on Vector Models and Indexing
The heterogeneity and high update frequency of hematologic oncology R&D documents pose specific requirements for vector model and index construction. First, for structured and semi-structured data in clinical reports, vector models must capture semantic associations, such as the link between a patient's specific gene mutation and a particular treatment plan. Second, gene sequencing data contains numerous short sequence snippets; vector models need to effectively process the semantic representation of short texts to avoid information loss. High update frequency necessitates that the indexing system supports efficient incremental updates, ensuring new data is included in the retrieval scope promptly. Furthermore, specialized medical terminology and units, such as CD34+ cell counts or FLT3-ITD mutations, require vector models to accurately understand their contextual meaning. This prevents ambiguity or insufficient generalization during vectorization, which would otherwise impact subsequent similarity matching accuracy.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
embedding_model | text-embedding-3-large or bge-large-zh-v1.5 | Enhances semantic understanding of biomedical terminology, particularly for fine distinctions in gene sequences and drug mechanisms of action. |
Chunk size (Chunk Length) | 500-800 characters | Balances context retention for long documents with semantic focus for short texts, especially applicable to structured tabular data within clinical trial reports. |
Chunk Overlap Length (Chunk Overlap Length) | 100-150 characters | Ensures contextual continuity across segments, crucial for texts involving multi-step treatment plans or complex pathological descriptions. |
Recall count (Recall Count) | Top 8-12 items | Considering the complexity and interconnectedness of hematologic oncology research, this increases recall to cover more potentially relevant information and improve recall rate. |
Similarity threshold (Similarity Threshold) | Calibrated by actual measurement | Calibrates against the similarity distribution of hematologic oncology-specific terms and concepts, ensuring highly relevant results are recalled and irrelevant information is avoided. |
Rerank result count (Reranked Return Count) | Top 5 items | Further refines initial recall results using a reranking model, prioritizing the most relevant clinical data or research conclusions for the query. |
Common Pitfalls
- Failing to re-index an existing knowledge base after changing the
embedding_model. This leads to a decline in query result quality because the new and old models generate vectors in different spaces, preventing accurate similarity matching with older vectors. - When importing documents containing extensive gene sequences or experimental data tables, the chunking strategy does not adequately consider the integrity of this special content. This results in fragmented chunks, losing critical gene loci or numerical associations, and impacting subsequent information extraction.
- During index construction, specific fields in clinical trial reports, such as
ECOGscores orCR/PR/SD/PD, are not pre-processed or tagged. This prevents the vector model from effectively distinguishing their semantic weight, leading to insufficient sensitivity to these important indicators during retrieval.
Validation Steps
- Select multiple test queries containing key terms (e.g.,
BCR-ABLfusion gene,CAR-Ttherapy). Check if the recalled results include directly or indirectly related documents containing these terms, and evaluate their ranking. - Upload a document containing structured tables (e.g., patient medication dosages and adverse reactions). After indexing, check if precise retrieval is possible using specific numerical values or units from the table, and verify context completeness.
- Compare query effectiveness using different
embedding_modelconfigurations. Through manual evaluation or expert review, determine which model better understands the specialized semantics of hematologic oncology and establish it as a baseline. - Regularly track changes in retrieval performance after new data imports, especially for new drug clinical trial data, to ensure the incremental update mechanism reflects the latest research progress in a timely manner.
Note: The values provided are common starting points. They should be measured against specific datasets and use cases.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.