Data Characteristics
Hematologic oncology R&D documents come from diverse sources. These include clinical trial reports, gene sequencing data, drug mechanism of action research papers, pathology analysis reports, and regulatory agency guidelines. Documents update frequently, especially with new drug development and clinical research advancements. Document structures are often semi-structured or unstructured text, such as PDF reports, Word documents, and specialized journal articles with many tables and figures. Fields and units involve complex medical terminology, gene loci, protein expression levels, cell counts, drug dosages (e.g., mg/kg), and biomarker concentrations (e.g., ng/mL). Specific medical units and abbreviations are common.
Constraints on Knowledge Base Retrieval and Recall
The data characteristics of hematologic oncology R&D documents impose specific requirements on knowledge base retrieval and recall. Semi-structured and unstructured data demand advanced text parsing to accurately identify and extract key information. Examples include disease subtypes, gene mutation types, drug targets, and clinical response data. High update frequency requires efficient incremental update mechanisms in the knowledge base to ensure timely retrieval results. Specialized terminology and complex units in documents mean simple keyword matching can lead to omissions or misunderstandings. This requires combining domain-specific dictionaries and semantic understanding for precise matching. Furthermore, data within tables and figures, if not effectively structured, will be difficult for the knowledge base to index, leading to missing recall of critical quantitative information. Therefore, the recall stage must integrate and associate multimodal information (text, tabular data).
Configuration Strategy
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Segment Length) | 800–1200 characters (characters) | Balances text context completeness with embedding model processing efficiency, suitable for long sentences and complex logic in medical documents. |
Chunk Overlap Length (Segment Overlap Length) | 100–150 characters (characters) | Ensures context continuity, preventing critical information from being cut at segment boundaries, suitable for documents dense with specialized terminology. |
Recall count (Number of Retrieved Items) | 8–12 entries (items) | Considering the complexity of hematologic oncology knowledge, this increases the number of retrieved items to cover more potentially relevant information. |
Similarity threshold (Similarity Threshold) | 0.75–0.85 | Domain knowledge requires high precision. A higher threshold reduces interference from irrelevant or generalized content. |
Rerank result count (Number of Reranked Items) | 3–5 entries (items) | After filtering by the reranking model, this provides the most relevant and refined results, reducing the processing burden on downstream models. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds (seconds) | Processing large clinical trial reports or research papers can take a long time, preventing timeouts. |
Common Pitfalls
- The error
worker terminated due to reaching memory liduring knowledge base creation often occurs when processing a single oversized file or when the total size of files uploaded in a batch is too large, leading to out-of-memory issues. - Missing critical gene loci or drug dosage information in retrieval results indicates that the document parsing stage failed to effectively identify and index structured data within tables or figures.
- In the knowledge base search module, incorrect assignment of
Knowledge Base Variablereferences leads to retrieval failures or inaccurate results. This typically stems from variable naming or reference paths not matching the actual configuration.
Verification Steps
- Select a typical hematologic oncology R&D document containing key clinical data and genetic information. Manually verify that critical fields (e.g., gene mutation types, drug dosages) are correctly extracted and indexed after document parsing.
- Use queries containing specific disease subtypes, drug names, and biomarkers. Test whether knowledge base retrieval results include highly relevant original document segments. Check if retrieved items cover the core content of the query intent.
- Simulate real-world application scenarios. Perform multiple retrievals via API or interface. Observe the precision and diversity of returned results under the configured
Similarity threshold(Similarity Threshold) andRerank result count(Number of Reranked Items). Adjust thresholds based on feedback from domain experts. - Check system logs to ensure no
PARSE_FILE_TIMEOUT_SECONDSrelated timeout errors or memory limit warnings occur during file uploads and knowledge base updates.
The values provided are common starting points. Measure them against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.