Data Characteristics in this Category
Infection control data sources are diverse. They include electronic medical records (EMR), laboratory test reports, microbial culture results, medication records, and hospital-acquired infection incident reports. Data updates frequently. For example, lab results may update hourly, and patient vital signs transmit in real-time. Document structures are typically semi-structured or unstructured. EMRs contain extensive free-text descriptions, while lab reports are more structured with defined test items and result fields. Fields and units are highly specific. Examples include microbial names, antimicrobial susceptibility profiles, minimum inhibitory concentration (MIC) values for antibiotics (typically in μg/mL), and ICD-10 codes for infection sites. These require precise identification and parsing.
Constraints Imposed by These Characteristics on Workflow Orchestration
High-frequency data updates require workflows to support near real-time data retrieval and processing. This prevents pre-screening with outdated information. Semi-structured and unstructured documents make traditional structured data processing difficult to apply directly. Stronger natural language processing (NLP) capabilities are necessary for information extraction and entity recognition. The presence of specific fields and units requires defining precise regular expressions or using domain knowledge graphs for mapping within the workflow. This ensures data parsing accuracy. For example, MIC value parsing must strictly differentiate units and thresholds for different antibiotics. Any deviation can lead to pre-screening misjudgments. Additionally, dispersed data sources require workflows to integrate multiple data interfaces and handle data format conversions between different systems.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
maxContext | 4000 characters | Accommodates the verbosity of medical texts, preventing truncation of key information. |
Recall count (Recall Count) | Top 10 | Increases coverage of relevant documents, improving pre-screening accuracy. |
Similarity threshold (Similarity Threshold) | 0.75 | Balances recall and precision, filtering out irrelevant information. |
Rerank result count (Rerank Return Count) | Top 5 | Reduces subsequent processing load, focusing on the most relevant content. |
PARSE_FILE_TIMEOUT_SECONDS | 300 seconds | Accommodates parsing time for large medical record documents. |
API_REQUEST_TIMEOUT_SECONDS | 60 seconds | Ensures timely responses for external laboratory system API calls. |
Three Common Mistakes
- The output variable
resultis empty. This occurs when the regular expressions in the information extraction step fail to match complex medical text patterns. - Knowledge base search results are too few. This occurs when the
Similarity threshold(Similarity Threshold) is set too high, filtering out relevant but not perfectly matching documents. - Workflow execution times out. This occurs when
PARSE_FILE_TIMEOUT_SECONDSorAPI_REQUEST_TIMEOUT_SECONDSare configured too low, not adequately accounting for large file parsing or external interface response delays.
How to Confirm Correct Configuration
- Simulate real case data. Check if each workflow execution accurately extracts key microbial names, MIC values, and infection site codes.
- Compare pre-screening results with expert manual judgments. Evaluate the sensitivity and specificity of the pre-screening and adjust the
Similarity threshold(Similarity Threshold) accordingly. - Monitor workflow logs. Check for timeout errors under
API_REQUEST_TIMEOUT_SECONDSandPARSE_FILE_TIMEOUT_SECONDSparameters. Ensure all data sources are processed in a timely manner.
The values provided are common starting points. Measure them against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.