Data Characteristics for This Domain
Protocols and SOP documents in the hematology-oncology field originate primarily from national health commissions, drug administration agencies, internal hospital regulations, and clinical guidelines published by professional societies. These documents are typically in PDF, Word, or scanned image formats. Content covers disease diagnostic criteria, treatment plans, drug usage specifications, clinical trial ethical guidelines, and adverse reaction handling procedures. Update frequency is relatively stable; national policies or guidelines are usually revised annually or biennially, while hospital internal SOPs are adjusted quarterly or semi-annually based on practical needs and the latest research. Document structures are rigorous, often organized into chapters, articles, and appendices, containing extensive specialized terminology, dosage units (e.g., mg/kg, U/L), time units (e.g., days, weeks, cycles), and descriptive fields for medical imaging reports.
Constraints Imposed by These Characteristics on Citation and Traceability
The rigor and specificity of hematology-oncology protocol documents demand precise citations to ensure authoritative answers. Their update frequency requires a knowledge base synchronization mechanism that balances timeliness and stability, preventing citations of outdated or superseded clauses. Complex document structures and specialized fields increase the difficulty of text segmentation and retrieval, potentially causing the model to lose critical information when understanding context. Numerical information, especially dosages and units, if cited incorrectly, directly impacts the accuracy of clinical decisions. Therefore, in the citation and traceability process, the integrity and accuracy of the original text must be ensured, with clear traceability to specific chapters and page numbers of the original document, to meet the high standards of verifiability required in the medical field.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Segment Length) | 300-500 characters (characters) | Balances paragraph completeness and retrieval efficiency, avoiding truncation of important medical terms or short sentences. |
Recall count (Recall Count) | Top 5-7 entries (top 5-7 entries) | Ensures sufficient contextual information, covering key content that may be dispersed across different paragraphs. |
Similarity threshold (Similarity Threshold) | 0.75-0.85 | Ensures highly relevant recall results to the query, filtering out unnecessary interfering information. |
Rerank result count (Reranked Return Count) | Top 3 entries (top 3 entries) | Highlights the most relevant citations, reduces model processing burden, and improves answer precision. |
maxContext | 3500 tokens | Accommodates more original text snippets, which is crucial for understanding complex diagnostic and treatment plans. |
Citation template (Citation Template) | “Original Text:{cited content} Source:{document name} {page number}” | Defines a clear citation format, facilitating traceability and verification, and meeting medical compliance requirements. |
Three Common Mistakes
- Model answers omit critical dosages or units in citations. This occurs when text segmentation separates numerical values from their units, leading to incomplete recall.
- The AI platform fails to retrieve relevant protocols from the knowledge base, indicating "no relevant information found." This happens when scanned PDF documents are uploaded directly without OCR processing, making the text content unindexable.
- Answer content shows semantic deviation from the original text in the knowledge base. This is due to a
Similarity threshold(similarity threshold) set too low, which retrieves many non-exact matching segments, interfering with the model's judgment.
How to Confirm Proper Configuration
- Select typical questions in the hematology-oncology domain. Verify whether the model's answers include key information and check if the cited
document nameandpage numberare accurate. - Randomly select protocol documents from the knowledge base and simulate user queries. Check if professional fields like dosages, units, and time in the answers completely match the original text.
- Review the knowledge base recall logs in the FastGPT backend. Analyze the actual recall effectiveness under the
Recall count(recall count) andSimilarity threshold(similarity threshold) to determine if all relevant segments are recalled. - For frequently updated protocols, regularly upload new document versions and query to verify if the model cites the latest version of the content.
Note: The values provided are common starting points. Measure performance against your own samples to determine optimal settings.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.