Data Characteristics
Infection control data sources typically include internal hospital infection control manuals, regulations, SOPs, disease control guidelines, normative documents from national health commissions, and historical infection event reports. These documents are often in PDF, Word, or scanned image formats, with varying degrees of structure. Update frequency varies: national and local policies may be revised annually or released irregularly, while internal hospital SOPs are updated 1–2 times per year based on policy adjustments or practical situations. Document content includes numerous professional terms, medical abbreviations, test indicators, and units such as "CFU/ml" (colony-forming units per milliliter), "%" (percentage), "hours," and "days." Fields may involve infection sites, pathogen types, antibiotic usage, disinfectant concentrations, and extensive unstructured descriptive text.
Constraints on Reference Sourcing and Traceability
The specialized nature and frequent updates of infection control documents require precise reference sourcing, down to specific clauses or paragraphs, to ensure information timeliness. Diverse document formats, especially scanned images, demand high OCR recognition accuracy and text extraction capabilities, directly impacting RAG recall accuracy. The abundance of medical terminology and unstructured descriptions can challenge general word segmentation and semantic matching models, necessitating targeted optimization for improved relevance ranking. Analyzing historical infection event reports often requires tracing back to specific timestamps and event details, which mandates reference mechanisms supporting chronological traceability. Any reference error or omission can lead to incorrect infection control advice, directly affecting patient safety and medical quality.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Segment Length) | 800–1200 characters | Ensures a single segment contains complete medical concepts or operational steps, preventing context fragmentation. |
Recall count (Recall Count) | Top 5 | Balances recall breadth with computational resource consumption, covering core relevant information. |
Similarity threshold (Similarity Threshold) | 0.75–0.85 | Addresses the need for precise matching of medical terminology, improving recall accuracy. |
Rerank result count (Rerank Return Count) | Top 3 | Streamlines the final presented references while maintaining quality, improving information digestion efficiency. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Provides sufficient file parsing time for large PDFs or scanned documents, preventing timeout interruptions. |
OCR_ENABLED | True | Ensures scanned or image-based infection control documents are correctly recognized and text extracted. |
Common Misconfigurations
- Reference results contain numerous irrelevant or low-relevance paragraphs. This occurs when the
Similarity threshold(Similarity Threshold) is set too low, leading to an overly broad RAG recall range. - AI responses fail to cite critical regulatory clauses, instead referencing general descriptions. This may stem from an improper
Chunk size(Segment Length) setting, causing key information to be truncated or mixed with non-critical information. - When users query specific historical infection events, the system cannot provide corresponding document references, or the referenced content lacks important details. This is typically due to poor OCR recognition quality or document parsing timeouts, preventing relevant information from being effectively indexed.
Verification Steps
- Select 10 infection control documents in different formats (PDF, Word, scanned images). Upload them and observe their indexing status to confirm all documents are successfully parsed without errors.
- Conduct at least 20 simulated queries on core infection control issues. Check the reference sources in the AI responses to confirm they point to specific documents, page numbers, or paragraphs, and that the content is highly relevant to the question.
- Randomly select 5 cited paragraphs from AI responses. Compare them against the original documents to verify the accuracy and completeness of the cited content and check for any missing critical information.
The values provided are common starting points. Measure them against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.