Data Characteristics
Solid tumor clinical trial data primarily originates from the National Medical Products Administration (NMPA) Center for Drug Evaluation (CDE), the U.S. National Institutes of Health (NIH) ClinicalTrials.gov database, and institutional review board (IRB) public disclosures. This data exists in a mix of structured and unstructured formats, with update frequencies typically weekly or monthly. Document structures include trial protocols, informed consent forms, and case report form (CRF) templates. Core fields cover trial name, investigational drug, indication, inclusion criteria, exclusion criteria, primary endpoints, secondary endpoints, study centers, and study phase. Some fields, such as inclusion and exclusion criteria, are often described in natural language, involving complex medical terminology and numerical ranges, for example, "tumor maximum diameter ≤ 5 cm" or "ECOG score 0-1".
Constraints Imposed by These Characteristics on Multi-turn Conversation and Prompt Engineering
Key information in solid tumor clinical trial data, such as inclusion/exclusion criteria and tumor staging, is often presented as unstructured text. This text contains extensive medical jargon, abbreviations, and numerical ranges. The multi-turn conversation system must accurately parse these complex descriptions during user queries and map them to structured or semi-structured data within the knowledge base. Multi-turn conversations require context understanding to continuously filter and match eligible trials as the user progressively refines their condition description. Prompt design must guide the model to focus on core elements like inclusion/exclusion criteria, disease progression status, and prior treatment history, avoiding generic responses. The data update frequency dictates the timeliness of knowledge base recall. Prompts should encourage the model to prioritize the latest data and, when necessary, indicate potential information lag.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Segment Length) | 800–1200 characters | Inclusion/exclusion criteria descriptions for solid tumor clinical trials are often long. This length ensures a complete semantic unit within a single segment, preventing crucial information from being truncated. |
Recall count (Recall Count) | 8–12 entries | This balances matching accuracy with model processing capability, recalling more potentially relevant trial protocols to improve recall rate. |
Similarity threshold (Similarity Threshold) | 0.75 | The field of solid tumors demands high precision in terminology. A threshold that is too low introduces excessive noise, while one that is too high might miss valid information. |
Rerank result count (Reranked Return Count) | 5 entries | Reranked models can more accurately filter the trials that best meet user needs, reducing the user's screening burden. |
maxContext | 4096 tokens | This ensures sufficient user input history and knowledge base recall content can be accommodated in multi-turn conversations, maintaining conversational coherence. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Clinical trial protocol files are typically large, requiring a longer parsing time to avoid file processing failures due to timeouts. |
Three Common Mistakes
- The model fails to cite or paraphrase
xlsxfile content from the knowledge base during a conversation. This occurs if the knowledge base was not correctly configured for file parsing and embedding when the application was created, preventing the model from accessing or understanding its internal data. - When the user inputs "PD-L1 expression positive," the model fails to correctly match relevant trials. This happens if the prompt does not sufficiently guide the model to associate natural language descriptions with standardized values in the
PD-L1 expression statusfield within the knowledge base. - In multi-turn conversations, the model repeatedly asks for information already provided, leading to a poor user experience. This is due to a
maxContextparameter set too low, preventing the model from effectively remembering previous conversation history and losing context.
How to Confirm Correct Configuration
- Upload a PDF or Word document containing multiple solid tumor clinical trial inclusion/exclusion criteria. Ensure the knowledge base correctly parses and embeds its content.
- Conduct multi-turn simulated conversations. Gradually refine patient characteristics (e.g., tumor type, staging, gene mutations). Observe if the model consistently filters for eligible trial lists and cites specific information from the knowledge base.
- Test extreme scenarios, such as inputting vague medical terms or abbreviations. Check if the model, guided by prompts, requests further clarification from the user or offers similar options.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.