Data Characteristics
Data for regulations and Standard Operating Procedures (SOPs) in hematologic oncology primarily originate from clinical practice guidelines published by national health commissions, expert consensuses from professional societies, and internal rules and operating procedures from medical institutions. These documents are typically in PDF, Word, or scanned image formats. Content covers disease diagnostic criteria, treatment plans, drug usage specifications, nursing procedures, and ethical review requirements. Update frequency is relatively stable; national guidelines are usually revised every 3-5 years, while institutional SOPs may be updated annually based on practical needs. Document structures are rigorous, often using chapters, sub-sections, and numbered items. They contain extensive medical terminology, abbreviations, dosage units (e.g., mg/kg, IU), time units (e.g., h, min), and normal ranges for laboratory indicators.
Constraints Imposed by These Characteristics on Model Access and Configuration
The rigorous structure and high density of specialized terminology in hematologic oncology regulation documents require models to effectively identify chapter boundaries during text segmentation to avoid semantic fragmentation. The presence of numerous medical abbreviations and dosage units demands higher semantic understanding precision from vectorization models. Models need domain knowledge to differentiate subtle nuances between similar terms. The relatively fixed document update cycle means that model training and knowledge base construction require periodic full or incremental updates to ensure information timeliness. Furthermore, scanned image documents rely on high-quality OCR recognition to guarantee text content integrity and accuracy. Any recognition error can directly impact the reliability of model responses.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Segment Length) | 800–1200 characters | Ensures each segment contains complete medical concepts or operational steps, preventing semantic fragmentation. |
Chunk Overlap Length (Segment Overlap Length) | 100–200 characters | Retains contextual information, enhances semantic coherence across paragraphs, and aids in understanding process details. |
Recall count (Recall Count) | Top 8 | Hematologic oncology regulations often involve multiple aspects; increasing recall covers more relevant provisions. |
Similarity threshold (Similarity Threshold) | 0.75 | Guarantees the precision of recalled content, avoiding interference from irrelevant or low-relevance clauses. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Provides ample time for OCR recognition and text parsing when processing large PDFs or scanned documents. |
UPLOAD_FILE_MAX_SIZE | 100 MB | Accommodates the upload of large regulatory documents containing numerous images or complex layouts. |
Three Common Mistakes
- After uploading large PDFs or scanned documents, the system remains unresponsive for an extended period or displays "parsing failed." This occurs because the
PARSE_FILE_TIMEOUT_SECONDSparameter is set too low, not providing enough processing time for the OCR engine. - The model provides inaccurate answers to questions involving dosages or laboratory indicators, manifesting as incorrect values or units. This is due to the vector model's insufficient understanding of specialized medical abbreviations and units, or text segmentation separating units from values.
- When asked about a specific regulatory clause, the model fails to return relevant content or returns incomplete items. This may stem from an inappropriate
Chunk size(Segment Length) setting, causing critical information to be split across different segments, or an insufficientRecall count(Recall Count) failing to cover all relevant context.
How to Verify Configuration
- Upload a hematologic oncology diagnostic and treatment guideline containing complex charts and extensive medical terminology. Check if it is successfully parsed. Verify the completeness of key terminology and process descriptions through knowledge base search.
- Ask multiple open-ended questions about the diagnostic criteria or treatment plans for a specific disease (e.g., Acute Myeloid Leukemia). Verify whether the model can accurately extract and integrate relevant clauses from the knowledge base.
- Select clauses from the regulations involving dosage calculations or time points. Ask questions and compare the numerical values and units returned by the model against the original text. Verify its understanding of upstream and downstream processes.
- Simulate a real clinical scenario by asking questions with ambiguous keywords. Observe whether the model's recall results cover potentially relevant regulatory clauses and check the filtering effectiveness of the
Similarity threshold(Similarity Threshold).
The values given are common starting points and should be measured against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.