Data Characteristics
Pharmacovigilance data originates from clinical trial reports, real-world study data, post-marketing adverse event reports (ADRs), drug labels, and regulatory guidelines. These documents update frequently, especially drug labels and regulatory guidelines, which revise based on new research or adverse event monitoring. Document structures are complex and diverse. They include structured tabular data (e.g., patient information, drug information, event descriptions in ADR reports) and extensive unstructured text (e.g., detailed clinical observation records, expert evaluation reports). Fields and units are specific. For example, drug dosages often use milligrams (mg), micrograms (μg), or units (U). Administration routes vary (e.g., oral, injection). Adverse event severity grading and causality assessment follow industry standard terminology and coding systems.
Constraints Imposed by Data Characteristics on Multi-turn Conversations and Prompts
The complex structure and specialized terminology of pharmacovigilance documents challenge the accuracy and depth of multi-turn conversations. Extracting key information from unstructured text requires prompts with strong semantic understanding and contextual relevance. For example, the onset time, duration, outcome, and suspected drug association of an adverse event often scatter across different paragraphs or even different documents. Multi-turn conversations must track this dispersed information and integrate and infer it in subsequent turns. Frequently updated data sources necessitate an efficient knowledge base update mechanism to ensure the conversation model references the latest information. Furthermore, recognizing and standardizing specialized fields and units requires prompt design to consider medical terminology dictionaries and unit conversion rules to avoid errors due to misinterpretation.
Configuration Guidelines
| Configuration Item | Suggested Value | Rationale |
|---|---|---|
maxContext | 2000–3000 Tokens | Pharmacovigilance documents have strong contextual relevance. A longer conversation history is needed to maintain semantic coherence. |
Chunk size (Segment Length) | 500–800 Characters | Balances semantic completeness and retrieval efficiency. Avoids redundancy in long segments and loss of context in short segments. |
Recall count (Recall Count) | Top 8–12 Entries | Ensures coverage of potentially relevant information. Addresses multi-faceted knowledge requirements in complex queries. |
Similarity threshold (Similarity Threshold) | 0.75–0.85 | Strictly controls the precision of recall results. Reduces interference from irrelevant or low-relevance documents. |
Rerank result count (Reranked Return Count) | 3–5 Entries | Focuses on the most relevant content. Improves the quality and efficiency of the final answer presented to the user. |
QUERY_REWRITE_MODEL | gpt-4o-mini or ERNIE-Speed | Enhances understanding of medical terminology and the ability to rewrite complex queries. Optimizes recall effectiveness. |
Three Common Pitfalls
- Tracking of individual patient information and adverse event report numbers fails in multi-turn conversations. This manifests as the inability to reference or incorrect referencing of specific cases mentioned in earlier conversation turns. This occurs when the
chatIdparameter is not correctly passed or the model fails to effectively bind entity recognition with the conversation context. - Key numerical fields, such as drug dosage and frequency, are incorrectly parsed or omitted from documents. This leads to inaccurate drug association judgments. This usually happens when prompts do not sufficiently guide the model to extract and convert specific formats (e.g., "twice daily, 50 mg each time") into structured data.
- When processing large amounts of unstructured clinical observation text, the model experiences "hallucinations" or incomplete information extraction. This manifests as generated answers deviating from the original text or omitting important details. This occurs when the
Chunk size(Segment Length) is too small, leading to truncated context, or theSimilarity threshold(Similarity Threshold) is too low, recalling excessive noise.
How to Verify Configuration
- For adverse event queries related to a specific drug, verify that multi-turn conversations accurately identify and track patient IDs and event descriptions in reports. Ensure the
chatIdmechanism functions correctly. - Test the model's ability to correctly extract and standardize drug usage information across different dosage units (mg, μg, U) and administration frequencies (once daily, bid, tid). Cross-check extracted values with the original text.
- Randomly select multiple complex clinical trial reports. Perform keyword queries and multi-turn follow-up questions. Check the model's accuracy in identifying key causal relationships, time sequences, and severity assessments within unstructured text. Ensure recall and reranking results effectively support the conversation.
Note: The values provided are common starting points. Measure them against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.