Data Characteristics in This Domain
Bispecific antibody clinical trial data primarily originates from clinical trial registries (e.g., ClinicalTrials.gov, European Clinical Trials Register), public reports from pharmaceutical companies, academic journal papers, and regulatory agency databases. Data update frequencies vary; registries might update weekly, while academic papers update as they are published. Document structures typically include trial protocols, investigator brochures, patient recruitment criteria, and efficacy and safety data.
Specific data fields include target combinations (e.g., BCMAxCD3), antibody structure (e.g., BsAb type), administration routes (e.g., intravenous injection), dose escalation data, and specific biomarkers (e.g., PD-L1 expression levels). Units cover common dosage units (mg/kg), time units (weeks, months), and biological indicator units (ng/mL, %).
Constraints Imposed by These Characteristics on Multiturn Conversation and Prompt Design
The diverse sources and varying update frequencies of bispecific antibody data require a multiturn conversation system to integrate information from different sources during knowledge retrieval and assess information timeliness. Complex target combinations and antibody structure descriptions necessitate prompts that can accurately identify and parse these specialized terms, avoiding semantic confusion.
Detailed patient recruitment criteria, such as ECOG scores and specific gene mutation statuses, require multiturn conversations to guide users in progressively refining query conditions. Key fields like dosage and administration route in efficacy and safety data must be accurately extracted and logically evaluated within the conversation to support pre-screening decisions. Additionally, for non-numerical but critical screening conditions like biomarkers, prompt design must accommodate diverse expressions.
Configuration Settings
| Configuration Item | Suggested Value | Rationale for This Value |
|---|---|---|
Chunk size (Chunk Length) | 500 characters (500 characters) | Accommodates long sentences and paragraphs in clinical trial protocols, maintaining contextual integrity. |
Recall count (Retrieval Count) | Top 8 entries (Top 8 entries) | Ensures coverage of multiple relevant clinical studies for complex queries, enhancing retrieval diversity. |
Similarity threshold (Similarity Threshold) | 0.75 | Balances retrieval precision and breadth, filtering out low-relevance clinical trial information. |
maxContext | 6000 token | Accommodates multiturn conversation history, allowing users to progressively refine screening conditions and preventing loss of critical information. |
Rerank result count (Reranked Return Count) | Top 5 entries (Top 5 entries) | Optimizes output in later stages of multiturn conversations, focusing on the most relevant clinical trials to improve user experience. |
query_rewrite_model | gpt-4-turbo | Enhances understanding of complex medical terminology and ambiguous queries, improving query rewriting quality. |
Three Common Mistakes
- Symptom: The system fails to understand a user's detailed follow-up questions about specific target combinations (e.g.,
CD20xCD3) during a multiturn conversation. Reason: The prompt design does not effectively guide the model to identify and extract key entities within target combinations, or these entities are not managed as independent variables in the context. - Symptom: When a user inquires about safety data, the system returns side effect information inconsistent with the currently queried antibody type. Reason: During knowledge base chunking or vectorization, the strong correlation between safety data and specific antibody types is not adequately considered, leading to cross-contamination in retrieval results.
- Symptom: After multiple follow-up questions, the system still cannot provide clinical trials matching specific
ECOGscores or gene mutation statuses. Reason: Fields related to patient inclusion/exclusion criteria in the knowledge base are not effectively indexed, or the prompt design fails to translate these conditions into precise query instructions.
How to Confirm Correct Configuration
- Simulate multiturn pre-screening for typical bispecific antibodies (e.g.,
CD3xBCMA) to check if the system accurately identifies targets, indications, and key inclusion/exclusion criteria. - Input queries containing vague or colloquial descriptions and observe if the system can guide the user to refine conditions through multiturn conversation, ultimately providing relevant clinical trial information.
- For numerical (e.g., age range) and non-numerical (e.g., gene mutation status) conditions within patient inclusion/exclusion criteria, verify if the system can correctly filter and exclude non-conforming trials.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.