Multi-turn Conversation and Prompt Engineering for Solid Tumor Clinical Trial Pre-screening

Solid tumor clinical trial data primarily originates from the National Medical Products Administration (NMPA) Center for Drug Evaluation (CDE), the

Data Characteristics

Solid tumor clinical trial data primarily originates from the National Medical Products Administration (NMPA) Center for Drug Evaluation (CDE), the U.S. National Institutes of Health (NIH) ClinicalTrials.gov database, and institutional review board (IRB) public disclosures. This data exists in a mix of structured and unstructured formats, with update frequencies typically weekly or monthly. Document structures include trial protocols, informed consent forms, and case report form (CRF) templates. Core fields cover trial name, investigational drug, indication, inclusion criteria, exclusion criteria, primary endpoints, secondary endpoints, study centers, and study phase. Some fields, such as inclusion and exclusion criteria, are often described in natural language, involving complex medical terminology and numerical ranges, for example, "tumor maximum diameter ≤ 5 cm" or "ECOG score 0-1".

Constraints Imposed by These Characteristics on Multi-turn Conversation and Prompt Engineering

Key information in solid tumor clinical trial data, such as inclusion/exclusion criteria and tumor staging, is often presented as unstructured text. This text contains extensive medical jargon, abbreviations, and numerical ranges. The multi-turn conversation system must accurately parse these complex descriptions during user queries and map them to structured or semi-structured data within the knowledge base. Multi-turn conversations require context understanding to continuously filter and match eligible trials as the user progressively refines their condition description. Prompt design must guide the model to focus on core elements like inclusion/exclusion criteria, disease progression status, and prior treatment history, avoiding generic responses. The data update frequency dictates the timeliness of knowledge base recall. Prompts should encourage the model to prioritize the latest data and, when necessary, indicate potential information lag.

Configuration Settings

Configuration ItemRecommended ValueRationale
Chunk size (Segment Length)800–1200 charactersInclusion/exclusion criteria descriptions for solid tumor clinical trials are often long. This length ensures a complete semantic unit within a single segment, preventing crucial information from being truncated.
Recall count (Recall Count)8–12 entriesThis balances matching accuracy with model processing capability, recalling more potentially relevant trial protocols to improve recall rate.
Similarity threshold (Similarity Threshold)0.75The field of solid tumors demands high precision in terminology. A threshold that is too low introduces excessive noise, while one that is too high might miss valid information.
Rerank result count (Reranked Return Count)5 entriesReranked models can more accurately filter the trials that best meet user needs, reducing the user's screening burden.
maxContext4096 tokensThis ensures sufficient user input history and knowledge base recall content can be accommodated in multi-turn conversations, maintaining conversational coherence.
PARSE_FILE_TIMEOUT_SECONDS600 secondsClinical trial protocol files are typically large, requiring a longer parsing time to avoid file processing failures due to timeouts.

Three Common Mistakes

  • The model fails to cite or paraphrase xlsx file content from the knowledge base during a conversation. This occurs if the knowledge base was not correctly configured for file parsing and embedding when the application was created, preventing the model from accessing or understanding its internal data.
  • When the user inputs "PD-L1 expression positive," the model fails to correctly match relevant trials. This happens if the prompt does not sufficiently guide the model to associate natural language descriptions with standardized values in the PD-L1 expression status field within the knowledge base.
  • In multi-turn conversations, the model repeatedly asks for information already provided, leading to a poor user experience. This is due to a maxContext parameter set too low, preventing the model from effectively remembering previous conversation history and losing context.

How to Confirm Correct Configuration

  • Upload a PDF or Word document containing multiple solid tumor clinical trial inclusion/exclusion criteria. Ensure the knowledge base correctly parses and embeds its content.
  • Conduct multi-turn simulated conversations. Gradually refine patient characteristics (e.g., tumor type, staging, gene mutations). Observe if the model consistently filters for eligible trial lists and cites specific information from the knowledge base.
  • Test extreme scenarios, such as inputting vague medical terms or abbreviations. Check if the model, guided by prompts, requests further clarification from the user or offers similar options.

The values provided are common starting points and should be measured against the reader's own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.