Data Characteristics in This Domain
Hematologic oncology quality documents primarily originate from clinical trial protocols at various hospitals, drug development companies' submission materials, and review guidelines and technical requirements published by national drug regulatory agencies. These documents update frequently, especially clinical trial protocols and review guidelines, which may revise annually based on the latest research or policy adjustments. Document structures typically include numerous tables, graphs, flowcharts, and complex medical terminology. Examples include drug dosages, administration protocols, and adverse event monitoring indicators. Fields and units are highly specialized, such as mg/kg (milligrams/kilogram), µL (microliters), %CR (complete remission rate), and OS (overall survival). These often come with specific medical abbreviations and coding systems.
Constraints Imposed by These Characteristics on "Model Access and Configuration"
The specialized and complex nature of hematologic oncology quality documents places specific demands on model access and configuration. First, documents contain image and tabular data. The model requires multimodal processing capabilities to effectively parse and integrate this non-textual information into the knowledge base. Second, the prevalence of specialized terminology and abbreviations requires the model to accurately understand their semantics during vectorization and retrieval, preventing recall bias due to lexical ambiguity. High update frequency necessitates that the knowledge base supports efficient incremental update mechanisms and can differentiate between new and old document versions. Furthermore, precise identification of specific fields and units requires the model to combine preset rules or specialized Named Entity Recognition (NER) models during information extraction, improving the accuracy of structured information extraction.
Configuration Strategy
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
UPLOAD_FILE_MAX_SIZE | 500 MB | Hematologic oncology documents often contain high-resolution graphs and detailed tables, leading to large file sizes. This ensures complete uploads. |
Chunk size (Segment Length) | 800–1200 characters (characters) | Preserves the integrity of medical context, preventing truncation of critical information, while balancing retrieval efficiency. |
Recall count (Recall Count) | 8–12 entries (items) | Ensures coverage of sufficient relevant information snippets to address multi-faceted knowledge requirements in complex queries. |
Similarity threshold (Similarity Threshold) | Calibrate based on actual measurements | Adjusts precision and recall for specialized terms and abbreviations by evaluating retrieval results. |
ENABLE_IMAGE_PROCESSING | True | Hematologic oncology documents are rich in image information. Enabling image processing ensures graphs, flowcharts, etc., are parsed. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds (seconds) | Processing large PDF documents and multimodal content can be time-consuming. This prevents parsing failures due to timeouts. |
Three Common Mistakes
- Symptom: The model provides inaccurate answers to questions about medical graphs or flowcharts, or directly ignores image content. Reason: Image parsing functionality is not correctly enabled, or the model used does not integrate multimodal processing capabilities, preventing image information from being effectively converted into an understandable representation.
- Symptom: When querying specific drug dosages or detection indicators, the model returns results with incorrect units or values. Reason: The knowledge base failed to accurately identify and extract numerical entities with units during the document parsing phase, or did not assign sufficient semantic weight to these entities during vectorization.
- Symptom: A locally deployed large model cannot be called normally, and the API returns a
Connection refusederror. Reason: The local large model service port is not exposed externally, or theMODEL_API_URLconfigured in FastGPT does not match the correct address and port.
How to Confirm Correct Configuration
- Upload a hematologic oncology clinical trial protocol containing complex graphs and multi-page tables. Observe if its parsing status is successful and check if the knowledge base includes image-related text descriptions.
- Conduct question-and-answer tests for specific medical terms and abbreviations in the document (e.g.,
CAR-T,AML,PFS). Evaluate the model's understanding of these specialized terms and the accuracy of information recall. - Select a query containing specific values and units (e.g.,
20 mg/kg,CD34+Cell). Check if the model's answer correctly identifies and references these key data points and evaluate their contextual relevance.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.