Workflow Orchestration for Clinical Trial Pre-screening in Infection Control

Infection control data sources are diverse. They include electronic medical records (EMR), laboratory test reports, microbial culture results

Data Characteristics in this Category

Infection control data sources are diverse. They include electronic medical records (EMR), laboratory test reports, microbial culture results, medication records, and hospital-acquired infection incident reports. Data updates frequently. For example, lab results may update hourly, and patient vital signs transmit in real-time. Document structures are typically semi-structured or unstructured. EMRs contain extensive free-text descriptions, while lab reports are more structured with defined test items and result fields. Fields and units are highly specific. Examples include microbial names, antimicrobial susceptibility profiles, minimum inhibitory concentration (MIC) values for antibiotics (typically in μg/mL), and ICD-10 codes for infection sites. These require precise identification and parsing.

Constraints Imposed by These Characteristics on Workflow Orchestration

High-frequency data updates require workflows to support near real-time data retrieval and processing. This prevents pre-screening with outdated information. Semi-structured and unstructured documents make traditional structured data processing difficult to apply directly. Stronger natural language processing (NLP) capabilities are necessary for information extraction and entity recognition. The presence of specific fields and units requires defining precise regular expressions or using domain knowledge graphs for mapping within the workflow. This ensures data parsing accuracy. For example, MIC value parsing must strictly differentiate units and thresholds for different antibiotics. Any deviation can lead to pre-screening misjudgments. Additionally, dispersed data sources require workflows to integrate multiple data interfaces and handle data format conversions between different systems.

Configuration Guidelines

Configuration ItemRecommended ValueRationale
maxContext4000 charactersAccommodates the verbosity of medical texts, preventing truncation of key information.
Recall count (Recall Count)Top 10Increases coverage of relevant documents, improving pre-screening accuracy.
Similarity threshold (Similarity Threshold)0.75Balances recall and precision, filtering out irrelevant information.
Rerank result count (Rerank Return Count)Top 5Reduces subsequent processing load, focusing on the most relevant content.
PARSE_FILE_TIMEOUT_SECONDS300 secondsAccommodates parsing time for large medical record documents.
API_REQUEST_TIMEOUT_SECONDS60 secondsEnsures timely responses for external laboratory system API calls.

Three Common Mistakes

  • The output variable result is empty. This occurs when the regular expressions in the information extraction step fail to match complex medical text patterns.
  • Knowledge base search results are too few. This occurs when the Similarity threshold (Similarity Threshold) is set too high, filtering out relevant but not perfectly matching documents.
  • Workflow execution times out. This occurs when PARSE_FILE_TIMEOUT_SECONDS or API_REQUEST_TIMEOUT_SECONDS are configured too low, not adequately accounting for large file parsing or external interface response delays.

How to Confirm Correct Configuration

  • Simulate real case data. Check if each workflow execution accurately extracts key microbial names, MIC values, and infection site codes.
  • Compare pre-screening results with expert manual judgments. Evaluate the sensitivity and specificity of the pre-screening and adjust the Similarity threshold (Similarity Threshold) accordingly.
  • Monitor workflow logs. Check for timeout errors under API_REQUEST_TIMEOUT_SECONDS and PARSE_FILE_TIMEOUT_SECONDS parameters. Ensure all data sources are processed in a timely manner.

The values provided are common starting points. Measure them against your own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.