Data Characteristics in This Category
Infectious disease clinical trial pre-screening involves diverse data sources. These primarily include Electronic Health Record (EHR) systems, Laboratory Information Systems (LIS), Picture Archiving and Communication Systems (PACS), and Clinical Trial Management Systems (CTMS). Data update frequency varies by source. EHR and LIS data may update in real-time. CTMS trial protocols and patient recruitment information update in batches or phases. Document structures differ. EHR and LIS often contain unstructured clinical notes, structured lab results, and medication records. PACS data consists of medical image files. CTMS primarily uses structured tabular data, recording inclusion/exclusion criteria and patient progress. Fields and units are specific. Microbiological culture results often include strain names and antimicrobial susceptibility results (e.g., MIC values in μg/mL). Complete blood counts include white blood cell counts (units 10^9/L). Imaging reports contain lesion descriptions and measurements (units mm).
Constraints Imposed by These Characteristics on "Workflow Orchestration"
The high heterogeneity and multi-modality of infectious disease data impose specific requirements on workflow orchestration. Real-time EHR and LIS data require support for streaming or high-frequency batch processing to ensure timely pre-screening results. Unstructured text and image data need advanced Natural Language Processing (NLP) and Computer Vision (CV) components. These components convert the data into structured features for rule matching or model inference. Examples include extracting symptom descriptions from clinical notes or identifying lesion features from imaging reports. Specific fields and units in microbiological culture and antimicrobial susceptibility results require precise matching and parsing by the knowledge base. The rule engine in the workflow must handle complex logic. For instance, it combines patient history, lab results, and imaging findings to determine compliance with specific inclusion/exclusion criteria. Integrating and cleaning multi-source heterogeneous data is a critical prerequisite for pre-screening accuracy. This requires dedicated data preprocessing nodes.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
maxContext | 2048 tokens | Clinical notes and imaging report texts are often long, requiring a sufficient context window for processing. |
Chunk size (Segment Length) | 800–1200 characters | Ensures each segment contains complete semantics, preventing truncation of key information. |
Recall count (Recall Count) | Top 10 entries (Top 10 items) | Infectious disease inclusion/exclusion criteria often involve multiple aspects. Increasing recall covers more relevant knowledge. |
Similarity threshold (Similarity Threshold) | Calibrated by actual measurement, e.g., 0.75 | Balances recall and accuracy, avoiding interference from irrelevant information. |
Rerank result count (Rerank Return Count) | Top 5 entries (Top 5 items) | After reranking, focuses on the most relevant core information, reducing the model's processing burden. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Parsing time for large imaging reports or complex clinical documents. |
Three Common Mistakes
- AI chat node output is inaccurate. The prompt's definition of the citation content template is vague. This prevents the model from effectively using specialized terms and context recalled from the knowledge base.
- Knowledge base search node filtering results are not as expected. The knowledge base configuration does not correctly set user authentication variables. This prevents the node from filtering accessible knowledge bases based on user roles or permissions.
- The AI model option in the workflow is empty. The model service is not correctly registered or configured. This prevents the platform from recognizing available models.
How to Confirm Correct Configuration
- Validate the output of each node in the workflow. Ensure data flow is as expected. For example, check that the
MICvalue unit for microbiological culture results is correct. - Simulate various patient data, including edge cases and abnormal data. Observe if pre-screening results align with clinical expert judgment. Record false positives and false negatives.
- Check the knowledge base search node. Adjust
Similarity threshold(Similarity Threshold) andRecall count(Recall Count). Ensure accurate recall of key inclusion/exclusion criteria and disease characteristics. - Check the AI chat node. Ensure it provides reasonable and professional answers to typical infectious disease clinical questions based on knowledge base content.
The values provided are common starting points. Measure them against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.