Data Characteristics
Clinical decision support systems use quality documents. These documents originate from internal medical institution guidelines, clinical pathways, disease management manuals, drug inserts, medical ethics guidelines, and regulatory files. Document updates are relatively stable, typically quarterly or annually, and involve new drugs, updated treatment techniques, or policy changes. Documents are structured hierarchically, often in Word, PDF, or Markdown formats. They contain extensive medical terminology, abbreviations, and dosage units. Fields include disease names, symptom descriptions, diagnostic criteria, treatment plans, drug dosages, and precautions. Dosage units like mg/kg, IU, and ml require high accuracy.
Constraints on Knowledge Base Retrieval and Recall
The specialized and structured nature of quality documents challenges knowledge base chunking granularity and semantic understanding. Phrases and terms in clinical guidelines have high contextual dependency; simple text chunking can break semantic integrity. The precision required for critical information like drug dosages demands highly accurate recall results to avoid serious misinformation. Although update frequency is not high, each update can involve core clinical logic adjustments. Therefore, incremental updates and version management capabilities are crucial for the knowledge base. Additionally, cross-references and chart information in documents require recall mechanisms to handle multimodal or hyperlink associations.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Chunk Length) | 300–500 characters | Balances semantic completeness and information density per chunk, preventing critical information truncation. |
Chunk Overlap Length (Overlap Length) | 50–100 characters | Ensures contextual continuity between adjacent chunks, improving recall accuracy. |
Recall count (Recall Count) | Top 5 entries | Covers common questions, balancing recall breadth with subsequent processing efficiency. |
Similarity threshold (Similarity Threshold) | Calibrate with actual tests | Ensures high relevance of recall results to user queries, avoiding irrelevant information. |
Rerank result count (Rerank Count) | 3 entries | Further refines results, prioritizing the most accurate clinical recommendations. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Handles parsing requirements for large PDF or Word documents. |
Common Pitfalls
- Recall results contain many irrelevant clinical pathways or drug information. This occurs when the
Similarity threshold(Similarity Threshold) is set too low, failing to filter noise effectively. - The system fails to provide complete diagnostic or treatment plans for complex disease inquiries. This occurs when
Chunk size(Chunk Length) is too short, fragmenting critical information and affecting contextual completeness. - Newly published clinical guidelines are not accurately retrievable. This occurs when the knowledge base does not receive timely incremental updates, leading to outdated data.
Verification of Configuration
- Select multiple typical disease cases. Simulate user queries and check if recall results cover core diagnostic criteria and treatment plans.
- For drug queries containing precise dosage units, verify that dosage data in recall content matches original units like
mg/kgorIU. - After a knowledge base update, compare recall differences before and after the update. Confirm new content is effectively retrievable and verify the stability of historical query results.
- Examine log output. Focus on
recall_countandrerank_countfields to ensure recall and reranking quantities meet expectations.
The values provided are common starting points and should be measured against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.