Data Characteristics
Data for clinical decision support in pharmacovigilance comes from drug inserts, medical literature, case reports, clinical trial data, adverse drug reaction databases (e.g., FDA Adverse Event Reporting System, FAERS), and electronic health records (EHR). Update frequencies vary: drug inserts and FAERS data typically update quarterly or annually, while medical literature updates more frequently. Document structures are diverse. Drug inserts are often semi-structured text with standard fields like indications, contraindications, dosage, and adverse reactions. Case reports and medical literature are unstructured text. Field units include milligrams (mg) and grams (g) for drug dosage, and "once daily" or "twice weekly" for frequency. Adverse reaction descriptions are natural language text.
Constraints from these Characteristics on Multi-turn Conversations and Prompts
Data source diversity requires multi-turn dialogue systems to integrate information from various formats. This includes extracting numerical values from structured drug dosage information and understanding semantics from unstructured adverse reaction descriptions. Varying update frequencies mean the knowledge base must regularly synchronize with the latest data to ensure timely decision support. Complex document structures require prompt design to effectively guide the model in locating key information within vast texts and identifying relationships between different fields. For example, when querying adverse reactions for a specific drug, the system must differentiate between drug names, dosages, and adverse reaction events. Furthermore, the natural language nature of adverse reaction descriptions demands higher semantic understanding from multi-turn dialogues. The system must handle synonyms, abbreviations, and ambiguous statements, and clarify user intent through multiple interactions.
Configuration Settings
| Configuration Item | Suggested Value | Rationale |
|---|---|---|
maxContext | 8000 tokens | Accommodates long medical literature and multi-turn conversation history, preventing context loss. |
Chunk size | 500 characters | Adapts to longer paragraphs in medical texts, ensuring semantic completeness. |
Recall count | Top 10 entries | Increases the probability of retrieving relevant information from a vast knowledge base. |
Similarity threshold | 0.75 | Balances recall accuracy and recall rate, preventing interference from irrelevant information. |
Rerank result count | Top 5 entries | Optimizes the relevance and conciseness of the final presented results. |
temperature | 0.3 | Ensures the rigor and accuracy of model output, reducing hallucinations. |
Common Pitfalls
- Output content is raw Markdown format, not rendered: The prompt did not explicitly request rendered output, causing the model to return raw Markdown syntax strings.
- AI dialogue fails to accurately retrieve relevant information: The knowledge base does not sufficiently cover relevant medical terminology or adverse drug reaction cases, or the segmentation strategy splits key information.
- Failure to extract time information in the prompt: The prompt does not provide clear time range definitions or examples, making it difficult for the model to correctly identify and extract relative time concepts like "this year" or "this month."
Verification
- For different drug and adverse reaction queries, verify the accuracy and completeness of the decision support information returned by the system against drug inserts or professional literature.
- In multi-turn dialogues, test the system's ability to correctly understand the evolution of user intent and provide coherent and relevant responses based on previous conversation history.
- Check knowledge base synchronization logs to confirm that data sources like drug inserts and FAERS are updated at the expected frequency, and that updated data indexes are effective.
Note: The values provided are common starting points and should be measured against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.