Data Characteristics
Infection control data originates from Hospital Information Systems (HIS), Laboratory Information Systems (LIS), Electronic Medical Records (EMR), and various infection surveillance reports. Data updates frequently, typically hourly or daily. This is especially true for microbial culture results, antibiotic susceptibility test data, and patient vital signs. Document structures vary. They include unstructured physician notes, semi-structured nursing record templates, and structured lab reports and medication records. Fields include Patient ID, Hospitalization Number, Infection Site, Pathogen Name, Antibiotic Type, Dosage, Frequency, Treatment Start Date, End Date, and Resistance Results. Units include common medical measurements such as mg, ml, ℃, and times/day.
Constraints on Knowledge Base Retrieval and Recall
The high-frequency updates of infection control data require the knowledge base to support efficient incremental indexing. This ensures the timeliness of retrieval results. The complexity of document structures, especially the large volume of unstructured text, challenges the robustness of text vectorization models. Models must accurately extract key information from colloquialisms, abbreviations, and medical terminology. Multiple, heterogeneous data sources require refined data cleaning and standardization. This unifies field names and unit representations across different systems. This avoids semantic gaps during retrieval. For example, the same antibiotic may have aliases or abbreviations in different systems. Furthermore, the rigor of clinical trial pre-screening demands high accuracy and interpretability from retrieval results. This prevents misjudgments of patient eligibility due to inaccurate recall.
Configuration Settings
| Configuration Item | Suggested Value | Rationale |
|---|---|---|
Chunk size (Segment Length) | 300–500 characters | Balances contextual completeness with vectorization model processing efficiency. Avoids diluting key information in overly long segments. |
Recall count (Recall Count) | Top 8–12 items | Ensures sufficient candidate documents for re-ranking while controlling computational resource consumption. |
Similarity threshold (Similarity Threshold) | 0.75–0.85 | Infection control data requires high accuracy. A threshold that is too low introduces noise. A threshold that is too high may miss relevant information. |
Rerank result count (Re-ranked Return Count) | 3–5 items | Selects the most relevant documents after re-ranking, improving the precision of the final answer. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Allows sufficient parsing time for large clinical guidelines or complex case documents. |
maxContext | 3500 tokens | Accommodates the multi-level logic and detailed descriptions common in medical texts. Ensures the model receives enough context. |
Common Mistakes
- Retrieval results include treatment plans or medication information irrelevant to the patient. This occurs due to incomplete data cleaning, failing to remove irrelevant medical record templates or historical records.
- Key fields (e.g., pathogen name, drug resistance) are empty in recalled documents. This happens when the document parser fails to correctly identify and extract specific fields from semi-structured documents.
- System response time significantly increases, or requests time out. This is caused by setting
Recall count(Recall Count) too high, leading to excessive load on the vector database or an overloaded re-ranking model.
How to Verify Configuration
- Test with various typical query statements. Evaluate the coverage and accuracy of key information in the recall results.
- Review knowledge base logs. Check if the number of document segments and the
PARSE_FILE_TIMEOUT_SECONDSparameter meet expectations. - Select a batch of patient cases known to meet or not meet inclusion criteria. Input relevant information for pre-screening. Analyze how recalled documents support the judgment.
- Monitor system resource usage under different query pressures. Ensure that the
Recall count(Recall Count) andRerank result count(Re-ranked Return Count) configurations do not cause performance bottlenecks.
The values given are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.