Data Characteristics in This Domain
Pharmacovigilance data in health management originates from various sources: user-reported events, wearable device data, electronic health records from medical institutions, and post-market surveillance reports from pharmaceutical companies. Data update frequencies vary. User reports can be real-time, wearable device data is often continuous or synchronized at set intervals, while medical institution and pharmaceutical company reports follow fixed cycles. Document structures also differ. User reports are typically unstructured text with structured symptom checkboxes. Device data is mostly time-series numerical data. Medical records and pharmaceutical company reports include patient demographics, medication history, adverse event descriptions, and diagnostic results, often with many fields and medical abbreviations. Units are diverse, including dosage (mg, µg), time (hours, days), and physiological indicators (mmHg, mmol/L).
Constraints Imposed by These Characteristics on "Multi-turn Conversations and Prompts"
The heterogeneous nature and varying update frequencies of health management pharmacovigilance data place specific demands on multi-turn conversation context management and prompt design. Unstructured text requires stronger semantic understanding to extract key information and integrate it into the conversation context for subsequent inquiries. Time-series data requires the conversation system to understand temporal trends, such as asking about "blood pressure trends over the last week." Due to broad and frequently updated data sources, the conversation system needs real-time or near real-time data querying capabilities to ensure conversations are based on the latest information. The presence of medical abbreviations and specialized units requires prompts to include instructions for terminology explanation or unit conversion to improve AI understanding accuracy. Additionally, users may mention multiple medications and their dosages within a conversation, requiring the system to accurately track and associate these entities across multiple turns to avoid information confusion.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
maxContext | 8 turns | In health management, users may need to describe symptoms and medication history in detail. 8 turns cover most scenarios and prevent loss of critical information. |
Chunk size (Segment Length) | 500-800 characters | Adverse drug reaction descriptions can be lengthy and contain many details. This length helps maintain context completeness. |
Recall count (Recall Count) | Top 8-12 | Sufficient relevant drug information, adverse reaction cases, or medication guidelines must be recalled from the knowledge base to support multi-faceted analysis. |
Similarity threshold (Similarity Threshold) | 0.75 | Pharmacovigilance demands high information accuracy. A higher threshold helps filter for more precise matches and reduces misjudgments. |
PARSE_FILE_TIMEOUT_SECONDS | 180 seconds | Users may upload reports with complex medical terminology or multiple pages. Extending parsing time ensures complete file processing. |
Rerank result count (Reranked Return Count) | Top 5 | After reranking, returning the top 5 most relevant pieces of information ensures quality without information overload. |
Three Common Mistakes
- After uploading a file in the chat window, the system displays "File parsing failed" or no response. This may be due to the uploaded file size exceeding the
UPLOAD_FILE_MAX_SIZElimit, or an unsupported file format preventing the parser from processing it. - The AI misunderstands medication dosages or frequencies mentioned by the user in multi-turn conversations. This occurs because prompts lack clear unit identification and entity association instructions, preventing the AI from correctly distinguishing dosage information for different medications.
- A workflow contains multiple AI conversation nodes, and the final output is redundant, displaying all intermediate conversation results. This happens when the workflow configuration does not explicitly specify to output only the result of the last conversation node, or a
returnnode is missing for result filtering.
How to Confirm Correct Configuration
- Upload simulated adverse drug reaction reports in various formats (e.g., PDF, TXT) and sizes. Check if files are parsed successfully and if extracted key information is accurate.
- Conduct multi-turn conversation tests, simulating users describing medication use and adverse reactions. Observe if the AI accurately identifies drug names, dosages, frequencies, and adverse reaction symptoms, and maintains conversational coherence.
- Design scenarios with multiple AI conversation nodes within a workflow. Test if the final output displays only the expected final result by checking the content of the
outputfield. - Test the system's understanding of medical abbreviations and specialized units. For example, input "BID" or "QID" to see if the AI correctly interprets them as "twice a day" or "four times a day."
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.