Data Characteristics for This Category
Clinical Decision Support (CDS) systems, during registration and declaration document preparation, primarily use multimodal medical literature, clinical guidelines, drug inserts, disease diagnosis and treatment standards, and approved medical device registration certificates. Data update frequencies vary. For example, drug inserts and clinical guidelines may update annually or based on significant research advancements, while medical literature is continuously generated. Document structures are diverse, including structured database entries, semi-structured PDF documents (e.g., clinical trial reports, approval documents), and unstructured text (e.g., expert consensus, academic papers). Field and unit specificities include medical terminology, measurement units (e.g., mg/kg, mmol/L), disease codes (ICD-10), drug codes (ATC), and specific clinical indicators (e.g., CTCAE grades). These require precise identification and processing to ensure information consistency and accuracy.
Constraints Imposed by These Characteristics on Knowledge Base Retrieval and Recall
The heterogeneous nature of CDS data sources challenges knowledge base construction, requiring support for parsing and importing various file formats. For example, PDF-formatted clinical trial reports require special handling for tables and figures to effectively extract key information. Irregular data updates necessitate incremental update and version management capabilities to ensure retrieved information is always current, preventing the use of outdated guidelines or drug information. The large volume of specialized terminology, abbreviations, and polysemous words in documents, along with potential synonyms or homonyms across different sources, demands high precision in text preprocessing and vectorization models. This directly impacts retrieval accuracy. Furthermore, the medical field requires extremely high information accuracy. The completeness and relevance of recall results are critical; any omission or misleading information can affect the quality of final decision support.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Chunk Size) | 500–800 characters | Clinical documents have strong contextual relevance. A moderate length balances information completeness and retrieval efficiency. |
Chunk Overlap Length (Overlap Length) | 100–150 characters | Ensures context is not lost at chunk boundaries, improving recall rate for cross-paragraph information. |
Recall count (Recall Count) | Top 8–12 items | Medical information is complex. Increasing the recall count covers more potentially relevant document snippets. |
Similarity threshold (Similarity Threshold) | 0.75–0.85 | Ensures high relevance of retrieval results and reduces interference from irrelevant information. Adjust based on actual data. |
Rerank result count (Reranked Return Count) | Top 5 items | Selects the most relevant items for comprehensive analysis by the language model, while maintaining broad recall. |
PARSE_FILE_TIMEOUT_SECONDS | 300 seconds | Processing large clinical trial reports or detailed guidelines requires longer file parsing times. |
Three Common Pitfalls
- Knowledge base retrieval results are empty. This may be due to unstable connections to external search services like SearXNG or unexpected return formats.
- Uploaded PDF documents are not correctly chunked or extracted in the knowledge base. This usually occurs because the PDF's internal structure is complex, containing many images or scanned pages, preventing text extraction tools from identifying valid text.
- Retrieval results contain a large amount of irrelevant information. This happens when the default similarity threshold is too low, failing to effectively filter out low-relevance document snippets.
How to Verify Correct Configuration
- After uploading typical documents (e.g., drug inserts, clinical trial reports), check if the knowledge base chunks are reasonable and if key information is correctly extracted.
- Perform searches for core medical concepts or query statements. Observe if the recalled document snippets contain key information relevant to the query intent and if the recall count meets expectations.
- Adjust the
Similarity threshold(Similarity Threshold) and use a set of test queries for comparison. Evaluate the relevance and precision of retrieval results until a balance is found. - Attempt to import documents in different formats (e.g., structured PDFs, unstructured guidelines). Confirm that file parsing has no errors and that content is correctly indexed by the knowledge base.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.