Data Characteristics for This Category
Clinical trial pre-screening in the biomedical field primarily sources medical record quality control data from Hospital Information Systems (HIS), Electronic Medical Record systems (EMR), and Clinical Trial Management Systems (CTMS). Data updates frequently, potentially hourly or even minute-by-minute during patient enrollment and visits. Document structures are complex, including unstructured physician diagnostic notes, structured laboratory test results, imaging reports, medication records, and progress notes. Fields are diverse, involving medical terminology, abbreviations, units of measurement (e.g., mmol/L, ng/mL, mmHg), and time formats (e.g., YYYY-MM-DD HH:MM). This data often has high privacy and sensitivity.
Constraints Imposed by These Characteristics on "Multi-turn Conversation and Prompts"
The complexity and high update frequency of medical record data require multi-turn dialogue systems to acquire the latest data in real-time or near real-time, and accurately parse various medical terms and units. The presence of unstructured text increases the difficulty of information extraction, requiring prompts to guide the model to focus on key information. Data sensitivity demands strict adherence to data anonymization and access control during processing. Multi-turn conversations need to track patient progress and treatment plan adjustments, making context management critical. Prompt design must consider how to efficiently filter key indicators that meet clinical trial inclusion/exclusion criteria from massive amounts of data, and handle synonyms and near-synonyms of medical terms.
Configuration Settings
| Configuration Item | Recommended Value | Rationale for This Value |
|---|---|---|
maxContext | 8192 token | Ensures accommodation of complete patient medical record fragments and multi-turn conversation history, preventing loss of critical information. |
Chunk size (Segment Length) | 500 characters (characters) | Balances semantic completeness and recall efficiency, ensuring each segment contains sufficient context. |
Recall count (Recall Count) | Top 10 entries (top 10 items) | Improves the recall rate of relevant medical record information, covering more potential inclusion/exclusion criteria details. |
Similarity threshold (Similarity Threshold) | 0.75–0.85 | Filters out irrelevant medical record fragments, retaining only content highly matched with pre-screening standards. |
Rerank result count (Rerank Return Count) | Top 3 entries (top 3 items) | Further refines recall results, improves model processing efficiency, and focuses on the most core medical evidence. |
WORKFLOW_MAX_RUN_TIMES | 500 | Accommodates multiple conditional judgments and information extractions that may be involved in complex pre-screening processes. |
Three Common Mistakes
- Output variable values show accumulation, where a variable includes the result of the previous variable. This may be due to improper node configuration in the workflow, causing variables to be accumulated instead of being correctly overwritten or cleared.
- Input guidance is enabled and a vocabulary is configured, but the expected questions are not seen in the conversation interface. This may be due to incorrect vocabulary format or incorrect association with the corresponding input guidance component.
- After uploading an XLSX file to the knowledge base, the model cannot reiterate its content. This may be because the knowledge base did not correctly parse the table file content, or did not process the table data as retrievable text during querying.
How to Confirm Correct Configuration
- Simulate multiple clinical trial pre-screening scenarios to check if the dialogue system can accurately identify and extract key inclusion/exclusion criteria from medical records, such as specific diagnoses, laboratory indicators, or medication records.
- Verify that the system can correctly maintain patient context information in multi-turn conversations, such as changes in examination results at different time points, and make logical judgments based on this.
- Check the prompt's parsing capability for various medical terms and abbreviations, confirming that the system can understand and process synonyms and near-synonyms, such as
hypertensionandHTN. - Compare the system's pre-screening conclusions with manual evaluation results, especially focusing on the accuracy of judgments for borderline cases, and adjust the
Similarity threshold(similarity threshold) based on inconsistencies.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.