Data Characteristics in This Category
Pharmacovigilance data originates primarily from clinical trial reports, real-world evidence (RWE), adverse drug reaction (ADR) reports, post-marketing studies, and drug package inserts. Data update frequencies vary. Clinical trial reports and post-marketing studies are typically released periodically. ADR reports are real-time. Document structures are diverse, including structured tabular data, semi-structured report texts, and unstructured free text. Fields include drug names, dosages, administration routes, adverse reaction descriptions (MedDRA codes), basic patient information, and reporting sources. Adverse reaction descriptions often use specific medical terminology and classification systems. Units cover dosage (mg, g, ml), frequency (times/day, daily), and time (hours, days, weeks).
Constraints Imposed by These Characteristics on "Reference and Traceability"
The diversity and complexity of pharmacovigilance data place specific demands on reference and traceability. Real-time ADR reports require the knowledge base to rapidly index and provide the latest information, ensuring reference timeliness. The mix of structured data and unstructured text means distinguishing data types during referencing and extracting original snippets from different formats. The presence of specialized terms like MedDRA codes requires the model to accurately match them when understanding and referencing, avoiding traceability errors due to semantic deviations. Patient privacy and data sensitivity also mandate that the referencing mechanism performs necessary anonymization or restrictions when displaying original text, ensuring compliance.
Configuration Settings
| Configuration Item | Recommended Approach | Rationale for This Approach |
|---|---|---|
Chunk size | 500–800 characters | In pharmacovigilance documents, adverse event descriptions or study conclusions typically appear in moderately sized paragraphs. Excessive length dilutes core information; insufficient length may lose context. |
Recall count | 8–12 entries | Pharmacovigilance reports can involve multiple aspects. Increasing the number of recalled items helps cover more comprehensive adverse event details, patient characteristics, or study designs. |
Similarity threshold | 0.75–0.85 | The pharmacovigilance domain requires high precision in matching professional terms and facts. A higher similarity threshold helps ensure the accuracy and relevance of recalled content. |
Rerank result count | 4–6 entries | Re-ranking selects a small number of the most relevant results. This effectively focuses on key adverse reactions, dose relationships, or patient characteristics, improving traceability efficiency. |
maxContext | 8000 tokens | Pharmacovigilance analysis often requires integrating multiple document snippets. A larger context window helps large models understand complex logical relationships and multi-factor influences, preventing information truncation. |
Vector Model | text-embedding-ada-002 or bge-large-zh | A vector model performing well in the medical domain is necessary to accurately capture the semantic similarity of professional terms and medical descriptions. |
Three Common Mistakes
- The answer provides only a conclusion without citing original snippets from the database. This occurs when the model fails to correctly identify and utilize the RAG mechanism to extract relevant text from external databases during answer generation.
- The citation limit for knowledge base search results is too low, leading to the non-recall of relevant but not directly matching reports. This happens when
Recall countorSimilarity thresholdsettings are overly conservative, failing to adequately cover the diverse expressions in the pharmacovigilance domain. - Cited sources are inaccurate, or the cited content has low relevance to the question. This occurs when the
Vector Modelfails to effectively understand the unique medical terminology and context of the pharmacovigilance domain, leading to deviations in similarity calculations.
How to Confirm Proper Configuration
- Select a test document containing a typical adverse event description. After asking a question, check if the cited original snippets in the answer accurately point to the relevant description in that document and verify if
Recall countmeets expectations. - For a report containing MedDRA codes, try asking questions about specific adverse reactions. Verify if the answer can cite the original text containing that code or related descriptions, and check if
Similarity thresholdeffectively filters high-quality results. - Test with pharmacovigilance documents from different sources (e.g., clinical trial reports and ADR reports). Verify if the
Chunk sizeandmaxContextconfigurations maintain good information completeness when the model references data of different structures and formats.
Note: The values provided are common starting points and should be measured against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.