Data Characteristics
Pharmacovigilance data in health management comes from patient health records, electronic medical record systems, wearable device data, and patient-reported adverse event information. This data is often semi-structured or unstructured. Examples include physician progress notes, patient-reported symptoms, and lab results. Update frequency varies: patient health records may update in real-time with appointments or biometric data uploads, while adverse event reports depend on reporting cycles after an event. Document structures are diverse, including free text, structured medical terminology (e.g., ICD-10 codes, SNOMED CT concepts), drug package insert content, and various PDF or image formats of examination reports. Fields and units include dosage (mg, μg), frequency (times/day), duration (days, weeks), and vital signs (blood pressure mmHg, heart rate bpm). These may have multiple representations or abbreviations.
Constraints on Reference and Traceability
The diverse and unstructured nature of health management pharmacovigilance data creates challenges for reference and traceability. Free text content requires finer text segmentation and semantic understanding to ensure reference relevance. Real-time or high-frequency data sources require the RAG system to index quickly and maintain reference timeliness. Diverse document formats necessitate robust document parsing capabilities to extract key information from PDFs and images into citable text segments. The complexity of medical terminology can increase the difficulty of similarity matching, requiring domain dictionaries or knowledge graphs. Health data involving personal privacy requires strict adherence to data security and anonymization guidelines during referencing. This ensures sensitive content is not disclosed while providing traceability. Inconsistent field and unit representations can affect the accurate identification and referencing of critical information like dosage and frequency.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size | 800–1200 characters | Balances semantic completeness of long texts with recall efficiency of short texts; accommodates longer free-text entries like medical records. |
Chunk Overlap Length | 100–200 characters | Ensures contextual continuity, preventing information loss when referencing across segments, especially in progress notes. |
Recall count | Top 5 entries | Balances query response speed with information completeness; pharmacovigilance requires ensuring no critical information is missed. |
Similarity threshold | 0.75–0.85 | High similarity requirements for domain-specific terminology to avoid irrelevant interference while ensuring sufficient recall breadth. |
Rerank result count | 3 entries | Focuses on the most relevant content, reduces user reading burden, and quickly identifies core adverse event information. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Accommodates parsing time for large PDF drug package inserts or lengthy medical record documents. |
Common Pitfalls
- Reply contains only reference links without specific text: This usually happens when
Similarity thresholdis too high orRecall countis too low. The model cannot find enough relevant content to generate a reply and only returns the most highly matched source links. - References point to irrelevant or outdated documents: This may result from delayed data index updates, or the document parser failing to correctly identify the publication date or version number of a document, leading to references to older drug package inserts.
- Unable to reference context from multi-turn conversations: This occurs due to incorrect
maxContextparameter configuration, or the semantic correlation of multi-turn conversations not being effectively passed to the RAG module. This limits the reference scope to the current turn.
Validation Steps
- Select typical adverse drug reaction cases. Input simulated queries. Check if the cited paragraphs in the reply accurately point to specific descriptions in relevant medical records, package inserts, or reports.
- Upload health management documents in various formats (PDF, text, image). Verify that the system correctly parses each document and generates citable segments.
- Simulate scenarios where patients report similar symptoms at different times. Check if the data cited by the system reflects the latest information and can differentiate records from different time points.
The values provided are common starting points. Measure them against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.