Data Characteristics for This Category
Clinical Decision Support Systems (CDSS) rely on core data from authoritative medical journals, clinical guidelines, drug instructions published by the National Medical Products Administration (NMPA) or the U.S. Food and Drug Administration (FDA), clinical trial reports, disease diagnosis and treatment norms, and drug adverse reaction databases. Data updates frequently, especially after new drug approvals, clinical guideline revisions, or adverse event monitoring results. Document structures are primarily unstructured text, such as academic papers in PDF format, guidelines in HTML or Word documents, and drug instructions in XML or PDF files. Fields and units are highly specialized, covering diagnostic criteria (e.g., disease code ICD-10), treatment plans (drug names, dosage units mg/kg, administration routes), efficacy evaluation indicators (e.g., imaging reports, laboratory test results mmol/L, U/L), and adverse reaction terms (MedDRA codes). Accuracy and standardization of medical terminology are critical.
Constraints on "Citation and Traceability" from These Characteristics
The highly specialized and multi-source nature of CDSS data places strict demands on citation accuracy and traceability. First, most documents are unstructured text. This can lead to inaccuracies with traditional keyword-based retrieval. More refined text segmentation and semantic understanding are necessary. Second, frequent data updates mean citations must reflect the latest versions. Otherwise, outdated or incorrect clinical advice may be cited. Third, the sensitivity of specialized fields and units requires precise citation of specific values, units, and context. This avoids decision-making errors due to ambiguity. For example, drug dosage or test result citations must include the complete value, unit, and reference range, not just a paragraph. Furthermore, cross-referencing data from different sources, such as a paper citing another guideline, requires tracing back to the original authoritative source to ensure complete chain traceability.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size | 500–800 characters | Balances contextual completeness with retrieval accuracy. Avoids overly long paragraphs introducing irrelevant information or overly short paragraphs losing critical medical context. |
Recall count | Top 10–15 entries | Increases the number of recalled items to cover more potentially valid information, considering the complexity and potential relevance of medical literature. |
Similarity threshold | 0.78–0.85 | Ensures semantic relevance of recalled content. Filters out low-quality or irrelevant medical text segments. |
Rerank result count | Top 5 entries | Further improves the ranking of the most relevant content through re-ranking algorithms after initial retrieval. Prioritizes the display of key clinical evidence. |
Reference Link Format | URL + page number | Allows users to directly navigate to the specific location in the original document. Facilitates manual verification and auditing. |
content extraction Mode | Mixed Structured and Unstructured | Combines precise extraction of structured data (e.g., tables) with semantic understanding of unstructured text (e.g., discursive paragraphs). |
Three Common Mistakes
- Generated citations do not match the user's question. This occurs when the
Similarity threshold(similarity threshold) in the knowledge base retrieval node is set too low, recalling many irrelevant document fragments. - Citation sources appear as "unknown" or are missing. This happens when metadata is not correctly parsed during document upload, failing to extract the
source_urlordocument_titlefields. - Subsequent questions in workflow tests do not receive knowledge base citations. This usually occurs because the global variable
datasetidis not correctly passed or referenced in subsequent nodes of the workflow, preventing the knowledge base retrieval node from specifying the target dataset.
How to Confirm Correct Configuration
- Ask typical clinical questions. Check if each citation points to a specific original document URL and page number, and if it successfully navigates.
- Randomly select 10 citations. Manually verify if the cited content semantically matches the corresponding paragraph in the original document, with no critical information missing.
- Test medical queries of varying complexity. Observe if the number of recalled citations remains stable within the configured
Rerank result count(re-ranked return count) and if the content is highly relevant.
Values provided are common starting points. Measure against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.