Data Characteristics in this Domain
Infection control data originates from hospitals, disease control centers, and research institutions. Data updates typically occur quarterly or annually. However, daily real-time monitoring data, such as infection case statistics and antibiotic usage, is also common. Documents come in various forms: national or local infection control guidelines, clinical guidelines, academic papers, internal infection control manuals, case reports, and laboratory test results.
These documents are complex. They contain normative clauses, flowcharts, clinical data, and microbiological indicators. Fields often include pathogen names, resistance profiles, infection sites, antibiotic types, dosages, treatment durations, patient comorbidities, and treatment outcomes. Data strictly adheres to medical units like mg/kg, CFU/mL, and %.
Constraints on Multi-turn Conversations and Prompts
The complex structure and specialized fields of infection control documents challenge the accuracy and robustness of multi-turn conversations. First, the dialogue system must accurately understand user queries about specific normative clauses or clinical pathways. This requires deep parsing of the document's hierarchical structure. Second, the prevalence of specialized terminology and units demands precise prompt design to prevent model deviations or confusion in generating responses. For example, when asked about "the sensitivity of a certain antibiotic to a specific pathogen," the system must locate relevant resistance profile information from vast datasets and present it with correct units. In multi-turn conversations, users may progressively refine query conditions, for instance, from "resistance of a certain strain" to "resistance genes of that strain to cephalosporins." This requires the system to possess context memory and reasoning capabilities to maintain conversational coherence.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
maxContext | 6 | Infection control queries often require tracing more historical dialogue to maintain contextual coherence. |
Chunk size (Segment Length) | 800-1200 characters (characters) | Infection control document paragraphs are typically long, containing detailed clinical descriptions and normative clauses. |
Similarity threshold (Similarity Threshold) | 0.75 | Ensures recalled document segments are highly relevant to professional queries, reducing interference from irrelevant information. |
Recall count (Recall Count) | Top 5-7 entries (top 5-7 items) | Infection control issues often involve multiple related knowledge points, requiring more recall items for comprehensive judgment. |
temperature | 0.3 | The infection control domain demands high accuracy and professionalism in answers, reducing creative output. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds (seconds) | Parsing large PDF or scanned guideline files can take a long time. |
Common Pitfalls
- Numerical or unit errors in dialogue. For example, misidentifying
mgasg. This occurs because prompts do not explicitly specify rules for extracting numerical values and units. - Loss of context in multi-turn conversations. The model's answers deviate from the topic after several follow-up questions. This happens when
maxContextis configured too low, failing to effectively retain historical dialogue information. - Inability to locate specific normative clauses. The system returns only general introductions instead of precise sections or paragraphs. This results from an unreasonable document segmentation strategy that does not adequately consider the document's hierarchical structure.
Verification Steps
- Conduct multi-turn dialogue tests for core infection control questions. Check if answers in each turn are accurate, coherent, and effectively utilize historical dialogue information.
- Randomly select infection control document segments containing numerical values and units. Ask related questions and verify if the system's output values and units match the original text.
- Test with infection control documents of different structures (e.g., tables, flowcharts, plain text). Verify if the system can accurately extract and structure information from them.
- For specific normative clauses, ask about their content and scope. Check if the system can precisely recall and cite the corresponding sections in the document.
The values provided are common starting points and should be measured against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.