Multiturn Conversation and Prompts for Antibody-Drug Conjugate (ADC) Clinical Trial Pre-screening

ADC clinical trial pre-screening data originates from clinical trial registries (e.g., ClinicalTrials.gov, CDE Clinical Trial Registration and

Data Characteristics

ADC clinical trial pre-screening data originates from clinical trial registries (e.g., ClinicalTrials.gov, CDE Clinical Trial Registration and Information Disclosure Platform), public R&D pipeline reports from biopharmaceutical companies, academic journal articles, patent documents, and internal research data. Data updates frequently, with new trial registrations, changes in patient recruitment status, and trial result publications continuously emerging. Document structures typically include protocols, patient informed consent forms (ICFs), investigator brochures (IBs), and case report forms (CRFs). These documents contain diverse fields such as drug names, targets, indications, inclusion/exclusion criteria, dosages, administration regimens, adverse events, and efficacy endpoints. They often involve specialized medical terminology, genomic data, biomarker expression levels, and complex statistical units (e.g., ng/mL, nM, AUC, Cmax, PFS, OS).

Constraints on "Multiturn Conversation and Prompts"

The specialized and complex nature of ADC clinical trial data places high demands on designing multiturn conversations and prompts. First, frequent data updates require efficient synchronization mechanisms for the knowledge base. This ensures real-time accuracy of conversation results and avoids recommendations based on outdated information. Second, diverse document structures require RAG retrieval models to effectively handle documents with varying formats and content depths, especially for inclusion/exclusion criteria that contain extensive logical judgments. The specialized nature of fields and specific units, such as target expression level thresholds, requires prompts to precisely guide the model to understand and extract critical numerical information, while avoiding confusion of similar terms or units. Additionally, multiturn conversations need to maintain contextual coherence. This supports users in iteratively querying and comparing different trial parameters, for example, comparing enrollment requirements for specific biomarker-positive patients across different ADC drugs.

Configuration Settings

Configuration ItemRecommended ValueRationale
Chunk size (Chunk Size)500–700 charactersInclusion/exclusion criteria and efficacy evaluations in clinical trial protocols are often logically dense paragraphs.
Recall count (Recall Count)8–12 itemsEnsures coverage of multiple relevant trials or different sections within the same trial.
Similarity threshold (Similarity Threshold)0.78Clinical terminology and expressions require high precision; a lower threshold risks introducing irrelevant information.
maxContext4096 tokensAllows carrying more historical conversation information, supporting complex multiturn comparisons.
Rerank result count (Rerank Return Count)5 itemsFilters the most relevant trial protocol snippets, reducing the model's processing load.
PARSE_FILE_TIMEOUT_SECONDS600 secondsHandles large PDF clinical trial protocols, ensuring complete parsing.

Common Pitfalls

  • Conversations that result in "no relevant trial information found" or "unable to determine if enrollment criteria are met" often indicate that the knowledge base index is not updated promptly or that the chunking strategy fails to effectively capture key information dispersed across different documents.
  • When users inquire about specific biomarker thresholds, if the model returns irrelevant values or incorrect units, this typically means the prompt did not clearly guide the model to focus on values and units, or document parsing failed to correctly identify unit-bearing numerical fields.
  • In multiturn conversations, if the model fails to remember previous constraints on targets or indications, leading to subsequent answers deviating from the topic, this occurs because maxContext is set too low, failing to retain sufficient conversational history context.

How to Verify Configuration

  • Ask multiturn questions about the latest ADC clinical trial protocols. Verify if the model accurately retrieves inclusion/exclusion criteria, primary endpoints, and secondary endpoints.
  • Randomly select at least 5 fields containing specialized numerical values and units (e.g., HER2 IHC 3+ or AUC 0-inf). Test if the model can correctly extract and interpret their meanings.
  • Simulate an engineer asking a series of iterative questions about a specific ADC drug, target, and patient subpopulation. Verify the coherence of the conversation context.
  • Upload a clinical trial summary report containing complex tables. Check if the model can accurately parse key data from the tables, such as adverse event rates for different dosage groups.

The values provided are common starting points. Measure them against your own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.