Data Characteristics for this Category
Respiratory clinical trial pre-screening data originates primarily from electronic health record (EHR) systems, patient-reported outcome (PRO) questionnaires, imaging reports (e.g., chest CT, X-ray), pulmonary function test results, and laboratory test data. Data update frequencies vary. EHRs and test results may update in real-time, while PRO questionnaires are filled out as needed. Document structures include semi-structured or unstructured text in EHRs, containing free-text fields for diagnoses, medical history, and medication history. Imaging reports are often descriptive text with structured conclusions. Field and unit examples include pulmonary function reports with FEV1 (forced expiratory volume in 1 second, in liters) and FVC (forced vital capacity, in liters). Laboratory test results involve various biomarkers (e.g., inflammatory factor CRP, in mg/L). This data is critical for screening criteria.
Constraints Imposed by These Features on Multiturn Conversation and Prompts
The complexity of respiratory diseases and the diversity of diagnostic data impose specific constraints on multiturn conversation and prompt design. Patient medical history and symptom descriptions are often open-ended text, making direct matching to structured screening criteria difficult. This requires the conversation system to understand unstructured text, especially when identifying key medical terms and exclusion criteria. Numerical indicators in imaging reports and pulmonary function data require precise extraction and comparison in the conversation. Prompt design must guide the model to focus on these numerical ranges and units. Multiturn conversations need to remember a patient's past medical history and medication use to avoid repetitive questioning and dynamically adjust subsequent questions based on context. Furthermore, since diagnostic criteria for respiratory diseases can involve multiple levels and mutually exclusive conditions, prompts must possess logical reasoning capabilities to ensure screening accuracy.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
maxContext | 800–1200 characters | Ensures coverage of patient's initial description, key medical history, and recent test results, preventing context loss. |
Recall Count | 5 entries | Balances recall precision and computational cost, focusing on core screening criteria documents. |
Similarity Threshold | 0.75 | Improves accuracy in matching relevant medical terms and symptom descriptions, reducing false positives. |
Segment Length | 400 characters | Optimizes the splitting of lengthy medical record text, ensuring each segment contains complete semantic information. |
Reranked Return Count | 3 entries | Further improves the ranking of the most relevant information based on initial recall. |
temperature | 0.3 | Maintains the rigor and objectivity of conversational responses, reducing generative hallucinations. |
Three Common Pitfalls
- The message "I did not select a knowledge base" appears in the conversation. This occurs because the knowledge base is not correctly associated with the current conversation, or the system fails to recognize a valid knowledge base ID.
- Certain key symptoms or test results provided by the patient are overlooked in the conversation, leading to inaccurate pre-screening results. This is typically because the prompt fails to effectively guide the model to extract specific medical entities or numerical values.
- The conversation repeatedly asks for the same information, resulting in a poor user experience. This is usually due to improper configuration or incorrect activation of the conversation
memorymechanism.
How to Confirm Correct Configuration
- Simulate multiple patient cases to verify whether the conversation system can accurately identify and extract key symptoms and test results related to respiratory diseases.
- Examine conversation logs to confirm that the system correctly remembers the patient's past medical history and previously provided information during multiturn interactions, avoiding repetitive questioning.
- Compare the preliminary screening conclusions provided by the system with manual judgments against predefined clinical trial screening criteria. Verify whether the main inclusion/exclusion criteria are applied correctly.
The values provided are common starting points and should be measured against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.