Data Characteristics
Pharmacovigilance data in hospital operations originates from electronic medical record systems, pharmacy management systems, adverse event reporting platforms, and internal clinical decision support systems. Data updates frequently. Adverse event reports are typically real-time or daily. Drug usage records change dynamically with patient visits and admissions/discharges. The document structure is primarily semi-structured and unstructured, including handwritten doctor's notes, nurse observation logs, patient self-reports, drug inserts, and medical imaging reports. Fields and units are highly specialized. For example, drug dosages are often in milligrams (mg), grams (g), or milliliters (ml). Dosing frequencies include once daily (qd) or twice daily (bid). Adverse reaction descriptions contain medical terminology and symptom descriptions.
Constraints from Data Characteristics on Multi-turn Conversations and Prompts
The high update frequency of pharmacovigilance data requires multi-turn conversation systems to quickly synchronize the latest information, avoiding recommendations based on outdated data. Semi-structured and unstructured document characteristics challenge RAG (Retrieval-Augmented Generation) in extracting key information, necessitating more refined text segmentation and embedding strategies. Specialized fields and units require prompt design to ensure the model accurately understands and generates medically compliant expressions, such as distinguishing magnitude differences between "mg" and "g". Colloquialisms or abbreviations in handwritten doctor's notes and patient self-reports increase the difficulty of entity recognition and intent understanding. Multi-turn conversations need stronger contextual understanding to correct or clarify user input, ensuring accurate information retrieval.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
maxContext | 6 | Covers common traceability depth for pharmacovigilance events |
Chunk size (Segment Length) | 800–1200 characters | Balances semantic completeness for lengthy medical record summaries and adverse event reports |
Recall count (Recall Count) | Top 10 entries | Ensures coverage of relevant information from diverse, heterogeneous data sources |
Similarity threshold (Similarity Threshold) | 0.75–0.85 | Filters highly relevant medical text, avoiding interference from irrelevant information |
Rerank result count (Rerank Return Count) | Top 5 entries | Focuses on the most critical pharmacovigilance-related evidence |
PARSE_FILE_TIMEOUT_SECONDS | 300 seconds | Accommodates parsing time for large medical record files or imaging reports |
Common Mistakes
- Files uploaded via the chat interface are not parsed, and no error message appears in the backend. This can occur if the file format is incompatible or the file size exceeds the
UPLOAD_FILE_MAX_SIZElimit, preventing the parsing task from starting. - The AI conversation module outputs multiple adverse drug reaction suggestions, but only the last one is desired. This might happen if multiple AI conversation nodes in the workflow configuration are set to output, and no clear filtering or aggregation rules are defined.
- The model's output for drug dosage or frequency is inconsistent with actual values. This could be due to the prompt not explicitly specifying unit conversion rules, or inconsistent records in the data source leading to model misinterpretation.
Validation Steps
- Simulate actual pharmacovigilance scenarios through multi-turn conversations. Submit medical record summaries and adverse event reports in various formats and content. Verify the model's accuracy in identifying key medical entities, dosage units, and dosing frequencies.
- Test uploading files of different sizes and formats. Observe if file parsing is normal, if parsing time is within
PARSE_FILE_TIMEOUT_SECONDS, and if the parsed text content is complete. - In complex multi-turn conversations, intentionally introduce ambiguity or incomplete information. Observe if the model can ask clarifying questions to obtain the necessary information. This assesses the reasonableness of the
maxContextsetting. - Compare the model's pharmacovigilance recommendations against real adverse reaction cases by inputting relevant medical record information. Assess compliance with clinical guidelines or experience to evaluate the effectiveness of
Similarity threshold(Similarity Threshold) andRecall count(Recall Count).
Note: The values provided are common starting points. Measure them against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.