Data Characteristics
Monoclonal antibody pharmacovigilance data comes from clinical trial reports, real-world studies, post-market surveillance databases (e.g., FDA Adverse Event Reporting System, FAERS; European Medicines Agency EudraVigilance), academic literature, and case reports. Data update frequencies vary; clinical trial data updates are relatively stable, while post-market surveillance data flows in continuously. Document structures are diverse, including structured case report forms, semi-structured medical texts (e.g., discharge summaries, follow-up records), and unstructured literature reviews. Fields and units are specific. For example, drug names often include batch information, dosages are in mg/kg or units, adverse event terms follow the MedDRA (Medical Dictionary for Regulatory Activities) coding system, event occurrence times are precise to the day, and severity grading follows CTCAE (Common Terminology Criteria for Adverse Events) standards.
Constraints on Multi-turn Conversations and Prompts
The diversity and specialized nature of monoclonal antibody pharmacovigilance data impose specific requirements on multi-turn conversations and prompt construction. First, the presence of professional terms like MedDRA codes requires prompt design to guide users in accurately describing adverse reactions and to support the system's recognition and association of these terms. Second, the precision of key fields such as dosage, time, and severity means multi-turn conversations must guide users to provide structured information to ensure data completeness and avoid ambiguity. For example, confusion over dosage units can lead to biased risk assessment. Third, the wide range of data sources requires prompts to cover different questioning scenarios. For instance, the questioning approach for clinical trial data should differ from that for post-market real-world data. Finally, the ability to parse unstructured medical text demands prompt design that can effectively extract key information and integrate it into the conversation flow, handling missing or inconsistent information.
Configuration Settings
| Configuration Item | Suggested Value | Rationale |
|---|---|---|
maxContext | 8 | Ensures the conversation covers a complete description of the adverse event, including key information such as drug, dosage, time, symptoms, and management measures. |
Chunk size | 512 characters | Accommodates longer descriptive paragraphs in medical texts, ensuring semantic completeness. |
Recall count | Top 10 entries | Increases the probability of retrieving relevant adverse event cases or drug information from the knowledge base, improving answer accuracy. |
Similarity threshold | 0.78 | Balances the precision and breadth of recall, avoiding interference from irrelevant information while not missing potential associations. |
Rerank result count | Top 5 entries | Further filters the most relevant items from the recall results, optimizing the quality of information presented to the user. |
queryRewrite | Enabled | Addresses colloquial or incomplete expressions users might use in multi-turn conversations, improving retrieval accuracy. |
Common Mistakes
- The system fails to correctly understand numerous medical terms in the conversation, leading to inaccurate replies or repeated clarification requests. This occurs because the knowledge base lacks sufficient embedding and indexing of relevant professional vocabulary.
- The user repeatedly provides adverse event occurrence times or dosages, but the system fails to integrate or extract effective information, leading to conversation breakdown or an ineffective loop. This occurs because the prompt design does not effectively guide users to provide structured information or parameter parsing rules are incomplete.
- The system suddenly loses context in a multi-turn conversation, requiring the user to repeat previous information. This occurs because the
maxContextparameter is set too low, unable to maintain a sufficiently long conversation history.
How to Verify Configuration
- Select 5-8 typical monoclonal antibody adverse event reporting scenarios. Simulate multi-turn conversations to evaluate the system's accuracy in extracting and understanding key information.
- Input descriptions containing MedDRA terms. Verify if the system correctly identifies and associates them with corresponding adverse event classifications in the knowledge base, and check the retrieved relevant cases.
- Intentionally omit or obscure adverse event dosage, time, and other information during a conversation. Check if the system can guide the user to complete necessary fields through multi-turn questioning, and observe the clarity and effectiveness of its prompts.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.