Multi-Turn Conversations and Prompts for Phase I Clinical Trial Prescreening

Phase I clinical trial prescreening data originates from subject recruitment questionnaires, medical examination reports, past medical history

Data Characteristics

Phase I clinical trial prescreening data originates from subject recruitment questionnaires, medical examination reports, past medical history records, and genetic test results. This data combines unstructured text (e.g., scanned handwritten doctor's notes, self-reported subject questionnaires) and structured data (e.g., laboratory indicators like complete blood count, urinalysis, liver and kidney function tests). Data updates frequently during the initial prescreening phase as subjects submit information, followed by additions and corrections. Document formats vary, including PDF informed consent forms, Word case report forms, and various image formats for medical imaging. Fields and units are highly specialized. For example, platelet count PLT uses the unit 10^9/L, and creatinine CRE uses umol/L. There are numerous medical abbreviations and specialized terms.

Constraints on Multi-Turn Conversations and Prompts

The highly heterogeneous nature of Phase I clinical trial prescreening data requires multi-turn dialogue systems to have robust document parsing and information extraction capabilities. Key information in unstructured text, such as past medical history and medication history, needs precise identification and structuring by large language models. The high frequency of data updates requires the knowledge base to quickly synchronize the latest information, preventing decisions based on outdated data. Highly specialized fields and units demand more precise prompt construction. Prompts must explicitly guide the large language model's understanding and processing of specific medical indicators. For example, when determining if a subject meets inclusion criteria, the system must accurately understand the normal ranges for AST and ALT. Furthermore, multi-turn conversations may involve interpreting medical imaging reports. This requires the system to have some multimodal processing potential, or at least the ability to guide users to upload images for manual assisted interpretation.

Configuration Settings

Configuration ItemRecommended ValueRationale
maxContext8Phase I clinical prescreening conversations are typically long, requiring extended context to trace medical and medication history.
Chunk size (Segment Length)800–1200 characters (characters)Medical text paragraphs are often long and contain detailed descriptions; a longer segment length helps maintain semantic integrity.
Recall count (Recall Count)Top 10 entries (Top 10)Ensures recall of sufficient relevant medical concepts and standards, improving decision accuracy.
Similarity threshold (Similarity Threshold)Calibrate based on actual measurementsBalances high recall and high accuracy, avoiding missed critical medical conditions or introduction of irrelevant information.
Rerank result count (Reranked Return Count)Top 5 entries (Top 5)Reranks recall results, prioritizing medical terms and standards most relevant to the current conversation intent.
temperature0.3Phase I clinical prescreening demands rigor and accuracy; a lower temperature value reduces model hallucination.

Common Pitfalls

  • Issue: The model returns an empty value or an error during a conversation, and the frontend displays a blank. Reason: The system does not effectively capture and handle empty returns from the large language model or API errors, leading to information disruption.
  • Issue: After a user query, the model does not offer "Did you mean" options. Reason: The knowledge base is not configured, is configured incorrectly, or the recalled content does not effectively trigger related questions.
  • Issue: The model misunderstands medical indicator units or normal ranges during a conversation. Reason: Prompts do not explicitly specify medical field units and reference values, or the knowledge base does not contain relevant standards.

How to Verify Configuration

  • Test different subject cases to verify the model's ability to correctly identify and extract key information such as medical history and medication history.
  • Simulate various consultation scenarios to check if the model maintains contextual coherence in multi-turn conversations and accurately cites knowledge base content.
  • Compare against real Phase I clinical trial inclusion/exclusion criteria to evaluate the model's judgment accuracy for specific medical indicators (e.g., ALT, eGFR).

The values provided are common starting points and should be measured against your own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.