Knowledge Base Retrieval and Recall for Infection Control Clinical Trial Pre-screening

Infection control data originates from Hospital Information Systems (HIS), Laboratory Information Systems (LIS), Electronic Medical Records (EMR), and

Data Characteristics

Infection control data originates from Hospital Information Systems (HIS), Laboratory Information Systems (LIS), Electronic Medical Records (EMR), and various infection surveillance reports. Data updates frequently, typically hourly or daily. This is especially true for microbial culture results, antibiotic susceptibility test data, and patient vital signs. Document structures vary. They include unstructured physician notes, semi-structured nursing record templates, and structured lab reports and medication records. Fields include Patient ID, Hospitalization Number, Infection Site, Pathogen Name, Antibiotic Type, Dosage, Frequency, Treatment Start Date, End Date, and Resistance Results. Units include common medical measurements such as mg, ml, ℃, and times/day.

Constraints on Knowledge Base Retrieval and Recall

The high-frequency updates of infection control data require the knowledge base to support efficient incremental indexing. This ensures the timeliness of retrieval results. The complexity of document structures, especially the large volume of unstructured text, challenges the robustness of text vectorization models. Models must accurately extract key information from colloquialisms, abbreviations, and medical terminology. Multiple, heterogeneous data sources require refined data cleaning and standardization. This unifies field names and unit representations across different systems. This avoids semantic gaps during retrieval. For example, the same antibiotic may have aliases or abbreviations in different systems. Furthermore, the rigor of clinical trial pre-screening demands high accuracy and interpretability from retrieval results. This prevents misjudgments of patient eligibility due to inaccurate recall.

Configuration Settings

Configuration ItemSuggested ValueRationale
Chunk size (Segment Length)300–500 charactersBalances contextual completeness with vectorization model processing efficiency. Avoids diluting key information in overly long segments.
Recall count (Recall Count)Top 8–12 itemsEnsures sufficient candidate documents for re-ranking while controlling computational resource consumption.
Similarity threshold (Similarity Threshold)0.75–0.85Infection control data requires high accuracy. A threshold that is too low introduces noise. A threshold that is too high may miss relevant information.
Rerank result count (Re-ranked Return Count)3–5 itemsSelects the most relevant documents after re-ranking, improving the precision of the final answer.
PARSE_FILE_TIMEOUT_SECONDS600 secondsAllows sufficient parsing time for large clinical guidelines or complex case documents.
maxContext3500 tokensAccommodates the multi-level logic and detailed descriptions common in medical texts. Ensures the model receives enough context.

Common Mistakes

  • Retrieval results include treatment plans or medication information irrelevant to the patient. This occurs due to incomplete data cleaning, failing to remove irrelevant medical record templates or historical records.
  • Key fields (e.g., pathogen name, drug resistance) are empty in recalled documents. This happens when the document parser fails to correctly identify and extract specific fields from semi-structured documents.
  • System response time significantly increases, or requests time out. This is caused by setting Recall count (Recall Count) too high, leading to excessive load on the vector database or an overloaded re-ranking model.

How to Verify Configuration

  • Test with various typical query statements. Evaluate the coverage and accuracy of key information in the recall results.
  • Review knowledge base logs. Check if the number of document segments and the PARSE_FILE_TIMEOUT_SECONDS parameter meet expectations.
  • Select a batch of patient cases known to meet or not meet inclusion criteria. Input relevant information for pre-screening. Analyze how recalled documents support the judgment.
  • Monitor system resource usage under different query pressures. Ensure that the Recall count (Recall Count) and Rerank result count (Re-ranked Return Count) configurations do not cause performance bottlenecks.

The values given are common starting points and should be measured against the reader's own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.