Data Characteristics in this Category
Infection control data primarily originates from Hospital Information Systems (HIS), Laboratory Information Systems (LIS), Electronic Medical Records (EMR), and microbiology test reports. This data includes patient demographics, diagnostic results, medication records, infection sites, pathogen types, antimicrobial resistance profiles, and treatment plans. Data update frequency is high. Some real-time data, such as vital signs and critical lab values, can update every minute. Progress notes and medication adjustments update hourly or daily. Document structures vary, including structured tabular data and extensive unstructured text like doctor's rounds notes, nursing records, and consultation opinions. Fields and units involve various medical terminologies, International Classification of Diseases (ICD-10) codes, drug dosage units (mg, IU), time units (hours, days), and qualitative descriptions or semi-quantitative indicators for microbiology culture results.
Constraints Imposed by these Characteristics on Multi-turn Conversation and Prompts
The high update frequency of infection control data requires the multi-turn conversation system to synchronize the latest information promptly. This prevents judgments based on outdated data. For example, patient medication adjustments or new microbiology test results can directly impact pre-screening decisions. Therefore, the knowledge base update mechanism must support high-frequency incremental synchronization. The prevalence of unstructured text demands more sophisticated prompt construction. Prompts need to effectively guide the model to extract key information from complex medical records, such as infection sites, pathogens, and resistance status. Diverse fields and units require prompts with strong semantic understanding capabilities to correctly parse medical terminology and values, avoiding misjudgments due to unit confusion. Additionally, multi-turn conversations may involve tracing historical medical courses. This requires the system to have robust context management capabilities to maintain a coherent understanding of the patient's overall condition across multiple interactions.
Configuration Settings
| Configuration Item | Recommended Value | Rationale for this Value |
|---|---|---|
maxContext | 8192 token | Needs to accommodate multi-turn conversation history, patient medical record summaries, and the current question to avoid information loss. |
Chunk size (Segment Length) | 500 characters (characters) | Balances RAG retrieval granularity and model processing efficiency, reducing information overload in a single segment. |
Recall count (Number of Retrieved Items) | Top 8 entries (top 8) | Increases coverage of relevant information and reduces the risk of missing critical data. |
Similarity threshold (Similarity Threshold) | 0.75 | Ensures strong relevance of retrieved content, reduces noise interference, and is calibrated through actual measurements. |
Rerank result count (Number of Reranked Items) | Top 3 entries (top 3) | Streamlines model input while maintaining relevance, improving inference efficiency. |
System Prompt | Include "As an Infection Control Expert" (As an infection control expert) | Defines the model's role, guiding it to analyze problems and provide recommendations from a professional perspective. |
Three Common Pitfalls
- The prompt "Knowledge base not selected" appears during a conversation. This occurs because the knowledge base configuration is not correctly linked to the current application, or the knowledge base query parameters are improperly set.
- The model's understanding of patient medical history is inconsistent across multi-turn conversations, exhibiting forgetfulness of key information mentioned in previous turns. This is due to insufficient
maxContextsettings for the context window or an improper memory management strategy. - Prompts fail to accurately extract key infection information, such as pathogens or resistance status, from unstructured medical records. This typically happens when the prompt structure is too simplistic and does not effectively guide the model for deep semantic extraction.
How to Confirm Proper Configuration
- Select multiple patient medical records containing typical infection control data. Simulate multi-turn conversations to verify if the model can accurately identify infection sites, pathogens, and resistance status.
- Randomly sample a proportion of conversation logs. Check the model's memory retention of critical patient information across multiple interactions to ensure contextual coherence.
- Test whether the model can timely reference and provide accurate advice from newly added infection control guidelines or drug information in the knowledge base during a conversation. This verifies the effectiveness of the knowledge base update mechanism.
The values provided are common starting points and should be measured against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.