Multi-Turn Conversations and Prompts for Mental Illness Clinical Trial Pre-screening

Mental illness clinical trial pre-screening data comes from various sources. These include Electronic Health Records (EHR), psychiatric assessment

Data Characteristics

Mental illness clinical trial pre-screening data comes from various sources. These include Electronic Health Records (EHR), psychiatric assessment scales, genomic data, neuroimaging reports, and medication history. This data often exists as unstructured text (e.g., doctor's diagnostic descriptions, patient-reported symptoms) and semi-structured data (e.g., scale scores). Data updates are relatively infrequent, typically every few weeks to months, coinciding with patient visits or assessment cycles. Document structures are complex. For example, an EHR may contain multiple sub-documents like outpatient records, inpatient records, and examination reports. Each sub-document further contains fields such as chief complaint, history of present illness, past medical history, and auxiliary examinations. Field names may have synonyms or abbreviations. Units are often score values, dosage units (mg), or frequencies (times/day).

Constraints Imposed by These Characteristics on Multi-Turn Conversations and Prompts

The complex document structure and high proportion of unstructured text in mental illness data challenge the information extraction and comprehension capabilities of multi-turn conversations. Low update frequency means less pressure on knowledge base timeliness, but the accuracy and consistency of historical data are crucial. Patient-reported symptoms and doctor's diagnostic descriptions often contain vague and subjective statements. This requires prompt design to focus more on context understanding and ambiguity resolution to prevent hallucinations or misjudgments by the model. Genomic and neuroimaging data are highly specialized, requiring the model to process specific domain terminology. Additionally, multi-turn conversations must accurately track changes in patient symptoms, medication history, and assessment results. This ensures the rigor of the pre-screening logic and prevents screening biases due to information omission or misinterpretation.

Configuration Guidelines

Configuration ItemRecommended ValueRationale for Recommendation
maxContext8192Mental illness clinical information is lengthy, requiring more context to support the coherence and accuracy of multi-turn conversations.
Chunk size500–700 charactersBalances semantic completeness and recall efficiency, avoiding excessive irrelevant information or truncation of critical information in long paragraphs.
Recall countTop 8–12 entriesEnsures coverage of key patient medical history, scale results, and genomic features, improving pre-screening accuracy.
Similarity threshold0.78–0.85Mental illness symptom descriptions often have subtle differences, requiring a higher threshold to ensure precise matching of recalled information.
Rerank result countTop 5 entriesRe-ranks recall results to prioritize the most relevant clinical information for the large model's judgment.
PROMPT_TEMPLATECalibrate by actual measurementCustomizes model reasoning for specific diagnostic criteria and exclusion conditions for mental illnesses.

Three Common Mistakes

  • Phenomenon: The model repeatedly asks for information already provided in the conversation. Reason: The maxContext parameter is set too low, preventing the model from remembering previous conversation turns or critical context.
  • Phenomenon: During pre-screening, the model recommends trials clearly unsuitable for the patient's condition. Reason: Specific contraindications or exclusion criteria for mental illnesses in the knowledge base were not effectively recalled, and the prompt failed to adequately emphasize this key information.
  • Phenomenon: Parts of an uploaded EHR document are not indexed. Reason: The PARSE_FILE_TIMEOUT_SECONDS parameter is too short. For complex documents containing a large amount of unstructured text, parsing times out, leading to partial content loss.

How to Verify Configuration

  • Conduct multiple simulated conversations covering typical patient symptoms, medication history, and genetic test results. Check if the model accurately understands and continuously tracks key information.
  • Validate knowledge base recall results. When querying complex medical records, observe if recalled items include all relevant diagnoses, assessment scales, and genetic markers. Check if Similarity threshold effectively filters out irrelevant information.
  • Upload various EHR documents with complex structures and varying lengths. Check if all critical information is correctly parsed and included in the knowledge base. Observe the impact of Chunk size and Rerank result count on the final reasoning results.

Note: The values provided are common starting points. Measure against your own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.