Data Characteristics
Quality documents in infectious diseases include clinical guidelines, diagnostic standards, treatment plans, drug instructions, pathogen detection report interpretations, and epidemiological surveillance data. Data sources are diverse, covering the World Health Organization (WHO), National Health Commission, Centers for Disease Control (CDC), academic journals, and drug regulatory databases. Updates are relatively frequent, especially with new infectious diseases or drug resistance variations; guidelines and plans may update within months. Document structures are primarily semi-structured and unstructured, containing extensive medical terminology, abbreviations, dosage units (e.g., mg/kg, IU), time units (e.g., hours, days, weeks), and specific disease classification codes (e.g., ICD-10).
Constraints Imposed by These Characteristics on Knowledge Base Retrieval and Recall
The rapid update frequency of infectious disease documents requires an efficient index update mechanism for the knowledge base. This ensures retrieval results are timely and avoids recalling outdated information. The intensive use of medical terminology and abbreviations in documents demands high semantic understanding from vector models. Models must accurately identify synonyms, near-synonyms, and hypernyms/hyponyms to improve retrieval accuracy. The semi-structured and unstructured nature of documents means keyword-based retrieval alone may miss critical information. This necessitates more refined text chunking and vector embedding techniques. Additionally, the mix of numerical information (like dosage and time) with descriptive text requires the knowledge base to handle multimodal information. Alternatively, preprocessing can convert numerical information into text features that vector models can effectively capture, supporting more precise retrieval filtering.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Chunk Size) | 500–800 characters | Balances contextual completeness with vector embedding efficiency. Avoids noise from overly long chunks and semantic loss from overly short chunks. |
Chunk Overlap Length (Chunk Overlap) | 50–100 characters | Ensures semantic continuity at chunk boundaries, preventing critical information from being split. |
Recall count (Recall Count) | 8–12 items | Balances recall breadth with the computational load for subsequent reranking or large language model processing. Ensures coverage of sufficient potentially relevant information. |
Similarity threshold (Similarity Threshold) | Calibrate by actual measurement | Adjust based on the distribution of the actual dataset and business tolerance, using a test set to ensure high recall and low false positive rates. |
Rerank result count (Reranked Return Count) | 3–5 items | Performs a secondary fine-grained ranking on recalled results, providing the most relevant core information and reducing the processing burden on large language models. |
embedding_model | text-embedding-ada-002 or bge-large-zh-v1.5 | These models offer good semantic understanding and vector representation for Chinese medical texts, effectively handling specialized terminology. |
Three Common Mistakes
- Retrieval results contain numerous irrelevant or outdated items. This occurs because the knowledge base index is not updated promptly, or the
Similarity threshold(Similarity Threshold) is set too low, leading to the recall of low-relevance documents. - Retrieval fails or is incomplete when dealing with specific disease names or drug abbreviations. The vector model's semantic understanding of specialized terminology is insufficient, or the text chunking strategy does not fully consider the integrity of professional vocabulary.
- After connecting the tool invocation module to the knowledge base, critical numerical information cannot be referenced from the retrieved results. The returned fields are empty or incorrectly formatted. This happens because the knowledge base failed to effectively parse and structure numerical data in documents during ingestion, or the invocation tool did not correctly configure the
json_pathextraction path.
How to Confirm Correct Configuration
- Select a batch of documents containing new pathogens or the latest treatment plans. Upload them to the knowledge base and immediately perform a retrieval to confirm the latest information is recalled. This assesses update timeliness.
- Execute a retrieval for a set of queries containing common medical abbreviations and complex medical terms. Check if the recalled results include semantically relevant, complete document snippets. Compare these against expert judgments of relevance. This assesses semantic understanding.
- Simulate real business scenarios by querying diagnostic standards or drug dosages for specific diseases. Check if the recalled results accurately extract key numerical information and verify its accuracy. This assesses the ability to extract structured information.
The values provided are common starting points. Measure them against specific samples to determine optimal settings.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.