Data Characteristics for this Category
Infectious disease quality documents include clinical guidelines, diagnostic and treatment protocols, disease prevention and control documents from national health authorities, drug inserts, laboratory testing standard operating procedures (SOPs), and case reports. These documents originate from various sources, including national authoritative bodies and internal medical institutions. Update cycles vary: clinical guidelines and protocols typically update every 2–5 years, but updates accelerate significantly for novel infectious diseases or drug resistance variations. Drug inserts are revised based on post-market surveillance and safety assessments. Document structures are typically hierarchical, with clear chapters, containing extensive medical terminology, dosage units, diagnostic criteria, and treatment processes. Common fields include disease names, pathogens, drug names, dosages (e.g., mg/kg), administration routes, treatment durations, test indicators (e.g., copies/mL), normal value ranges, and judgment criteria.
Constraints Imposed by These Characteristics on "Multi-Turn Conversations and Prompts"
The specialized and rigorous nature of infectious disease documents requires multi-turn conversations to precisely identify medical terminology when understanding user intent. The accuracy of critical numerical values like dosage and treatment duration directly impacts clinical practice. Therefore, prompt design must emphasize the extraction and validation of numbers and units. The periodic and sudden nature of document updates means the knowledge base needs flexible document version management to ensure conversations are based on the latest, most authoritative information. If a user's question involves a novel disease or rapidly evolving drug-resistant strain, the system must identify relevant updates in the knowledge base and guide the user to the latest information. Additionally, common hierarchical structures and cross-references in documents require the conversational system to effectively track context across multiple turns, avoid information fragmentation, and use prompts to guide the model in logical reasoning and information integration.
Configuration Settings
| Configuration Item | Recommended Value | Rationale for this Recommendation |
|---|---|---|
Chunk size (Segment Length) | 500–800 characters | Ensures each segment contains sufficient context while avoiding excessive length that could lead to information overload and reduced recall efficiency. |
Recall count (Number of Retrieved Items) | top 8 | Infectious disease documents are highly specialized, requiring more relevant paragraphs to support complex questions. Too many, however, introduce noise. |
Similarity threshold (Similarity Threshold) | 0.78–0.85 | Balances recall precision and breadth, preventing the retrieval of less relevant segments that could affect answer accuracy. |
Rerank result count (Number of Reranked Items) | top 5 | Further optimizes retrieval results by ranking the most relevant segments higher, improving model utilization efficiency. |
maxContext | 6000 tokens | Multi-turn conversations about infectious diseases often involve complex medical histories, test results, and treatment plans, requiring a longer context window. |
temperature | 0.2–0.4 | Ensures answers are rigorous and objective, reducing generative deviation and preventing speculative medical advice. |
Three Common Mistakes
- During a conversation, the model fails to identify treatment plans from the latest guideline version, instead citing outdated information. This occurs due to poor knowledge base document version management or prompts that do not explicitly prioritize the latest version.
- When a user asks about a specific drug dosage, the model returns a pure number without units or with incorrect units. This happens because prompts do not explicitly require the model to include units when extracting numerical values, and unit validation is not performed.
- A user's question involves cross-information from multiple document chapters, and the model only answers content from one chapter, leading to incomplete information. This is because prompts fail to effectively guide the model in cross-document or cross-chapter information integration and reasoning.
How to Confirm Proper Configuration
- Test with a series of questions containing both old and new version information to verify if the model consistently outputs key information from the latest document version.
- Prepare test questions involving numerical units like dosage and treatment duration. Check the accuracy and completeness of numerical units in the model's responses.
- Design complex questions requiring information integration across chapters or documents. Evaluate if the model maintains contextual coherence and provides integrated answers in multi-turn conversations.
- Monitor
input tokens quantityandoutput tokens quantityduring conversations to assess if the context window and retrieval strategy meet expectations, avoiding unnecessarytokensconsumption.
The values provided are common starting points. Measure them against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.