Data Characteristics in Infection Control Management
Infection control data primarily originates from Hospital Information Systems (HIS), Laboratory Information Systems (LIS), Electronic Medical Records (EMR), and specialized infection surveillance systems. This data updates frequently; for example, microbiology culture results, antimicrobial susceptibility reports, patient temperatures, and medication records may update in real-time or daily. Document structures are diverse, including unstructured physician order texts and nursing notes, semi-structured lab reports and imaging reports, and structured patient demographics and diagnostic codes. Fields include patient ID, ward, bed number, admission time, discharge time, diagnosis, surgical information, infection site, pathogen name, antimicrobial susceptibility results (e.g., MIC values, sensitive/resistant), antimicrobial drug name, dosage, administration route, and medication start/end times. Units typically include μg/mL for MIC values, mg, g, or IU for drug dosages, and dates and hours for time.
Constraints Imposed by These Characteristics on Tool Calling and Plugins
The high update frequency of infection control data requires tool calling to have near real-time data synchronization capabilities to avoid making decisions based on outdated information. The diversity of document structures necessitates preprocessing different data sources during tool calls, such as entity extraction from unstructured text and field mapping for structured data. Parsing semi-structured reports requires flexible template matching or large model-based extraction capabilities. Furthermore, sensitive information in infection control data (e.g., patient identity) demands high security and compliance for tool calling, requiring strict data anonymization and access control. The specificity of fields and units, such as MIC values in antimicrobial susceptibility results, requires tools to understand and correctly process these specialized terms and units to support accurate pharmacovigilance rule evaluation.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
UPLOAD_FILE_MAX_SIZE | 500 MB | Accommodates large electronic medical records and imaging uploads, balancing transfer efficiency and storage pressure. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Accounts for the parsing time of complex reports (e.g., pathology reports, imaging reports) to prevent task failures due to timeouts. |
maxContext | 8192 token | Ensures the ability to process long patient progress notes and multiple lab reports at once, maintaining context integrity. |
Chunk size | 800 characters | Balances semantic completeness and retrieval efficiency, adapting to the paragraph structure of medical texts. |
Similarity threshold | 0.75 | Improves the relevance of recall results, reducing interference from irrelevant medical literature or case snippets. |
Recall count | Top 10 entries | Ensures coverage of sufficient potentially relevant information while avoiding excessive redundant data that could affect subsequent processing. |
Rerank result count | Top 3 entries | Selects the most relevant entries from the initial recall to provide more focused evidence for decision-making. |
Common Pitfalls
- Key variables (e.g.,
patient_id,infection_site) are empty in subsequent steps after tool calling. This typically results from incorrect front-end data mapping or back-end parsing logic failing to extract required fields from raw data. - File upload errors with
Invalid file formatwhen calling a workflow via API. This usually occurs because theContent-Typeheader is not set correctly, or the uploaded file encoding does not match the API's expectations. - Knowledge base query results do not match expectations during custom RAG integration, failing to recall relevant infection control guidelines or drug instructions. This may be due to the knowledge base indexing strategy not adequately considering synonyms and hierarchical relationships of medical terms, leading to low retrieval accuracy.
Validation Steps
- Simulate real-world scenarios by uploading infection control documents in various formats (e.g., PDF lab reports, plain text physician orders). Check if the tool successfully parses and extracts predefined structured information, and observe logs for parsing failure messages.
- Execute workflows that include tool calls. Observe if the
tool_outputfield contains the expected results, and check if key parameters (e.g.,drug_name,dosage) have reasonable values and units. - For specific infection control cases, call the workflow via API with simulated data. Verify if the tool call results in the
workflow_outputalign with expected judgments (e.g., pharmacovigilance alerts, infection risk assessments) and if the output logic conforms to infection control guidelines. - Test the knowledge base retrieval function by inputting infection control-related medical terms or clinical questions. Verify if the
scorevalues of the recalled results are generally high and if the returned document snippets contain the core information of the query.
Note: The values provided are common starting points. Measure against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.