Data Characteristics
Infectious disease protocol data primarily comes from national health commission guidelines, clinical pathways, hospital infection control manuals, standard operating procedures (SOPs), and relevant regulations. These documents are typically in PDF, Word, or internal knowledge management system pages. Update cycles are relatively stable; national guidelines update every 3–5 years, while hospital SOPs may revise annually or adjust based on new epidemics. Document structures are often chapter-based, including introductions, definitions, diagnostic criteria, treatment plans, and prevention measures, with varying paragraph lengths. Fields include disease names, pathogens, drug names, dosages, administration routes, ICD-10 codes, microbial test result units (e.g., CFU/mL), and antibiotic sensitivity (e.g., MIC values).
Constraints on Knowledge Base Retrieval and Recall
The specialized and standardized nature of infectious disease protocol documents requires precise matching of medical terminology and specific codes, such as ICD-10. Documents often contain tables and figures; pure text parsing may miss critical information, affecting recall quality. Although update frequency is not high, any revision can impact core clinical decisions. The knowledge base must quickly identify and update affected knowledge points to avoid recalling outdated or incorrect protocols. Multi-chapter structures and cross-references between documents mean a single document's context may be incomplete, requiring cross-document associative recall. Furthermore, the precision of drug dosages and units demands high completeness and accuracy in retrieval results, preventing errors due to truncation or misidentification.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Segment Length) | 500–800 characters | Balances contextual completeness and retrieval efficiency, accommodating longer paragraphs common in SOPs. |
Chunk Overlap Length (Segment Overlap Length) | 100–150 characters | Ensures no information loss at paragraph boundaries, especially when discussing disease mechanisms across segments. |
Recall count (Recall Count) | 8–12 entries | Provides richer context, considering the complexity of infectious disease diagnosis and treatment. |
Similarity threshold (Similarity Threshold) | 0.75–0.85 | Ensures high relevance of retrieval results to the query, reducing interference from imprecise medical terms. |
Rerank result count (Reranked Return Count) | 3–5 entries | Focuses on the most critical and relevant protocol clauses, avoiding information overload. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Accommodates parsing time for large PDF or Word documents, preventing timeout errors. |
Common Pitfalls
- During search testing, entering professional terms may not recall expected results. This can happen if document parsing fails to correctly extract or index specific medical vocabulary, or if segmentation is too short, leading to loss of contextual semantics.
- After uploading a PDF file, query results may show "no relevant information found." This often occurs when the PDF is a scanned image without OCR processing, preventing the knowledge base from recognizing and indexing text content.
- Differences in question-answering effectiveness exist between API models and local models. Local models may show insufficient reference to knowledge base content, possibly due to limited parameters or fine-tuning data, resulting in lower understanding of specific domain knowledge compared to API models.
Verification Steps
- For core diseases (e.g., pneumonia, sepsis), query using different phrasing to observe if recall results include key diagnostic criteria and treatment plans.
- Check queries related to drug dosages and administration routes in the knowledge base. Verify that recalled text fully includes units and values.
- Search using unique
ICD-10codes or pathogen names from the documents. Confirm precise location of relevant protocol clauses. - Simulate a protocol update scenario. Upload a revised document and test if the knowledge base recalls the latest version of information, covering potential errors in previous versions.
The values provided are common starting points and should be measured against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.