Data Characteristics for this Category
Quality documentation data for Clinical Decision Support (CDS) systems primarily originates from authoritative medical guidelines, clinical pathways, disease diagnosis and treatment protocols, drug inserts, and medical literature. Data update frequency is relatively stable, typically quarterly or annually. However, urgent updates may occur with new drug approvals or guideline revisions. Document structures are mainly structured or semi-structured. They include critical information such as disease diagnostic criteria, treatment plans, drug contraindications, adverse reactions, dosage units (e.g., mg/kg, IU), and time units (e.g., hours, days). Documents often contain medical terminology, abbreviations, charts, and flowcharts. Fields frequently have strict medical definitions and standardized coding systems (e.g., ICD-10, SNOMED CT). Unit expressions are precise and diverse.
Constraints on "Model Access and Configuration" from these Characteristics
The data characteristics of CDS quality documentation impose specific requirements on model access and configuration. The rigorous structure and high density of specialized terminology demand strong semantic understanding from the model to avoid misjudgments due to insufficient medical background knowledge. Data update frequency dictates the knowledge base refresh strategy, requiring support for incremental updates and version management to ensure the timeliness and accuracy of decision support. Complex medical units and numerical ranges in documents require vector models to effectively distinguish and understand the context of numerical values during embedding, for example, the difference between mg/kg and mg. Furthermore, a large amount of semi-structured data requires more refined text segmentation and metadata extraction strategies to ensure that complete and relevant contextual information is provided during retrieval, preventing the loss of critical decision-making basis due to improper segmentation.
Configuration Settings
| Configuration Item | Suggested Value | Rationale |
|---|---|---|
Chunk size (Chunk Length) | 800–1200 characters | Ensures the completeness of medical terms and context, preventing critical information from being fragmented. |
Overlap Length | 100–200 characters | Ensures continuous context at chunk boundaries, improving retrieval relevance. |
Recall count (Retrieval Count) | Top 5–8 | Clinical decisions often require multi-faceted evidence. Increasing the retrieval count provides a more comprehensive reference. |
Similarity threshold (Similarity Threshold) | 0.75–0.85 | Clinical decisions demand high accuracy. A high threshold filters out irrelevant or weakly relevant document fragments. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Medical documents are often large and complex, requiring longer parsing times to avoid timeouts. |
maxContext | 3000–4000 tokens | Ensures the large language model can process longer medical contexts, preventing information truncation from affecting judgment. |
Common Pitfalls
- Vector access fails, with logs indicating "vector model cannot recognize input data type." This occurs due to an inappropriate vector model selection that does not support the specific encoding or format of medical text, or the model has not been fine-tuned for the medical domain.
- Application responses contain dosage or unit errors, such as misinterpreting
mgasg. This happens when numerical values and their associated units are not correctly identified and extracted during document parsing, leading to the loss of critical quantitative information during vector embedding. - Uploading large medical guideline files results in a "file parsing timeout" error. This is because the
PARSE_FILE_TIMEOUT_SECONDSparameter is set too low, not providing enough parsing time for complex documents.
Verification Steps
- Upload and parse various types of medical documents (e.g., guidelines, drug inserts). Check if parsing results are complete and if key fields and units in the documents are correctly identified.
- Ask multiple questions related to specific medical issues. Observe whether the document fragments retrieved by the model are accurate, relevant, and contain necessary quantitative information.
- Simulate an urgent update scenario by uploading new or revised documents. Verify if the model's response to the latest information is timely and accurate after the knowledge base update.
Note: The values provided are common starting points. Measure them against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.