Citation and Traceability for Clinical Decision Support Registration Document Preparation

Clinical Decision Support Systems (CDSS) rely on core data from authoritative medical journals, clinical guidelines, drug instructions published by

Data Characteristics for This Category

Clinical Decision Support Systems (CDSS) rely on core data from authoritative medical journals, clinical guidelines, drug instructions published by the National Medical Products Administration (NMPA) or the U.S. Food and Drug Administration (FDA), clinical trial reports, disease diagnosis and treatment norms, and drug adverse reaction databases. Data updates frequently, especially after new drug approvals, clinical guideline revisions, or adverse event monitoring results. Document structures are primarily unstructured text, such as academic papers in PDF format, guidelines in HTML or Word documents, and drug instructions in XML or PDF files. Fields and units are highly specialized, covering diagnostic criteria (e.g., disease code ICD-10), treatment plans (drug names, dosage units mg/kg, administration routes), efficacy evaluation indicators (e.g., imaging reports, laboratory test results mmol/L, U/L), and adverse reaction terms (MedDRA codes). Accuracy and standardization of medical terminology are critical.

Constraints on "Citation and Traceability" from These Characteristics

The highly specialized and multi-source nature of CDSS data places strict demands on citation accuracy and traceability. First, most documents are unstructured text. This can lead to inaccuracies with traditional keyword-based retrieval. More refined text segmentation and semantic understanding are necessary. Second, frequent data updates mean citations must reflect the latest versions. Otherwise, outdated or incorrect clinical advice may be cited. Third, the sensitivity of specialized fields and units requires precise citation of specific values, units, and context. This avoids decision-making errors due to ambiguity. For example, drug dosage or test result citations must include the complete value, unit, and reference range, not just a paragraph. Furthermore, cross-referencing data from different sources, such as a paper citing another guideline, requires tracing back to the original authoritative source to ensure complete chain traceability.

Configuration Settings

Configuration ItemRecommended ValueRationale
Chunk size500–800 charactersBalances contextual completeness with retrieval accuracy. Avoids overly long paragraphs introducing irrelevant information or overly short paragraphs losing critical medical context.
Recall countTop 10–15 entriesIncreases the number of recalled items to cover more potentially valid information, considering the complexity and potential relevance of medical literature.
Similarity threshold0.78–0.85Ensures semantic relevance of recalled content. Filters out low-quality or irrelevant medical text segments.
Rerank result countTop 5 entriesFurther improves the ranking of the most relevant content through re-ranking algorithms after initial retrieval. Prioritizes the display of key clinical evidence.
Reference Link FormatURL + page numberAllows users to directly navigate to the specific location in the original document. Facilitates manual verification and auditing.
content extraction ModeMixed Structured and UnstructuredCombines precise extraction of structured data (e.g., tables) with semantic understanding of unstructured text (e.g., discursive paragraphs).

Three Common Mistakes

  • Generated citations do not match the user's question. This occurs when the Similarity threshold (similarity threshold) in the knowledge base retrieval node is set too low, recalling many irrelevant document fragments.
  • Citation sources appear as "unknown" or are missing. This happens when metadata is not correctly parsed during document upload, failing to extract the source_url or document_title fields.
  • Subsequent questions in workflow tests do not receive knowledge base citations. This usually occurs because the global variable datasetid is not correctly passed or referenced in subsequent nodes of the workflow, preventing the knowledge base retrieval node from specifying the target dataset.

How to Confirm Correct Configuration

  • Ask typical clinical questions. Check if each citation points to a specific original document URL and page number, and if it successfully navigates.
  • Randomly select 10 citations. Manually verify if the cited content semantically matches the corresponding paragraph in the original document, with no critical information missing.
  • Test medical queries of varying complexity. Observe if the number of recalled citations remains stable within the configured Rerank result count (re-ranked return count) and if the content is highly relevant.

Values provided are common starting points. Measure against specific samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.