Data Characteristics
Pharmacovigilance data in retail pharmacies originates from pharmacy system records, customer medication feedback, pharmacist reports, and product information from drug suppliers. Data updates frequently. Daily sales data and medication feedback are generated in real-time. Drug instructions and contraindications typically update with new batches or are regularly released by regulatory bodies. Document structures vary, including structured sales records, unstructured customer verbal reports, semi-structured adverse reaction report forms, and standardized drug instructions. Fields and units are industry-specific, such as lot_number, expiry_date, dosage_unit (e.g., milligrams, tablets), administration_route, and free-text fields like adverse_event_description.
Constraints on Multiturn Conversation and Prompts
The high update frequency and diverse sources of retail pharmacy data demand real-time knowledge base updates for multiturn conversation systems. This ensures that drug information and pharmacovigilance data cited in conversations are the latest versions. Complex document structures, especially the presence of extensive free text and semi-structured reports, challenge accurate key information extraction, impacting conversation understanding and response accuracy. The strictness of fields like drug batch numbers and expiry dates requires the conversation system to match precisely during queries, avoiding misjudgments due to minor differences. Additionally, the colloquial, non-standardized descriptions in customer medication feedback test the robustness and semantic understanding capabilities of prompts. The system needs to identify potential adverse reaction signals from vague descriptions. This requires prompt design to balance standardized queries with natural language understanding.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
maxContext | 8 | Ensures coverage of typical medication consultations and adverse reaction reporting processes in multiturn conversations, preventing context loss and repetitive questioning. |
Chunk size (Segment Length) | 500-800 characters (characters) | Balances textual semantic integrity with recall efficiency, accommodating paragraph lengths in documents like drug instructions and adverse reaction reports. |
Recall count (Recall Count) | 7-10 entries (items) | Considering the specialized and associative nature of pharmacovigilance data, increasing the recall count helps cover more potentially relevant knowledge points. |
Similarity threshold (Similarity Threshold) | 0.78-0.85 | Reduces the false recall rate, especially when processing highly similar texts such as drug names and symptom descriptions, ensuring the accuracy of recall results. |
Rerank result count (Rerank Return Count) | 3-5 entries (items) | Performs a secondary filtering based on initial recall, focusing on the most relevant information to improve the accuracy and relevance of conversation responses. |
ENABLE_STREAM_OUTPUT | True | Enhances user experience, especially when querying complex drug information or adverse reaction reports, as users can see response progress in real-time. |
Common Pitfalls
- The system responds with "cannot find relevant drug information" or "cannot understand symptom description." This occurs when free text in the knowledge base is insufficiently preprocessed, failing to effectively extract and index key entities.
- Users repeatedly inquire about different batches or dosages of the same drug, but the system cannot differentiate and provide precise answers. This manifests as generic or repetitive information being returned. This happens due to a lack of effective differentiation and retrieval mechanisms for fine-grained attributes like drug batches and dosages in the knowledge base.
- When users describe adverse reactions, the system's response is inconsistent with the context or provides non-specific advice. This may be due to
maxContextbeing set too low, leading to truncation of conversation history and an inability to maintain a complete semantic chain.
Validation Steps
- Select a test case set covering various drugs, different batches, and typical adverse reaction scenarios. Execute multiturn conversation tests to verify the system's ability to accurately identify drugs and symptoms and provide relevant information.
- Randomly select drug instructions and adverse reaction reports from the knowledge base. Simulate user questions and check if the recalled knowledge snippets are complete and accurate. Compare them with original documents to confirm the effectiveness of
Similarity thresholdandRecall count. - Observe whether the system can extract key information from colloquial descriptions when processing free-text customer feedback and generate logically clear, professional responses. This evaluates the prompt's natural language understanding capability.
The values provided are common starting points and should be measured against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.