Pharmacovigilance Data Characteristics
Pharmacovigilance data originates from clinical trial reports, real-world evidence (RWE), adverse drug reaction (ADR) databases, medical literature, and drug labels. Data updates frequently. New drug approvals lead to a continuous influx of adverse reaction reports. Existing drug labels may also be revised due to new safety information. Document structures typically include structured fields and unstructured text, such as patient demographics, drug information, adverse event descriptions, diagnoses, and treatment measures. Fields and units are highly specialized. Examples include medical terms for adverse events (MedDRA codes), dosage units (mg, IU), time units (days, hours), and laboratory test results (mmol/L, U/L).
Constraints Imposed by Data Characteristics on Vector Models and Indexing
The high update frequency of pharmacovigilance data requires vector indexes to support efficient incremental updates. This ensures the timeliness of query results. Documents contain numerous specialized medical terms and abbreviations. This demands high semantic understanding from vector models; general models may struggle to accurately capture deep meanings. The mixed structured and unstructured document format means indexing strategies must support both field retrieval and semantic retrieval. For example, querying adverse reactions for a specific drug in a specific organ requires retrieving drug names from structured fields and descriptions of organs and adverse reactions from unstructured text. Key numerical information, such as time and dosage, requires specific vectorization or integration with metadata filtering to support precise queries.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk Size | 800–1200 characters | Adverse drug event descriptions often contain multiple pieces of information. This length effectively captures complete context while preventing excessively long chunks that introduce noise. |
Chunk Overlap | 100–200 characters | Ensures that related information across chunks is not lost, especially when describing event progression or causal relationships. |
Embedding Model | Specialized medical domain fine-tuned model | Improves the accuracy of understanding MedDRA terms, drug names, and clinical descriptions. |
Recall Count | Top 10–20 items | Pharmacovigilance queries typically require a broad initial recall range to cover potentially relevant information. |
Similarity Threshold | Calibrated by actual measurements | Determine based on the specific model and data distribution, using recall and precision curves to ensure highly relevant results. |
Reranked Output Count | 5–8 items | Balances user reading experience with information comprehensiveness, providing the most relevant few results. |
Common Pitfalls
- After uploading to the knowledge base, some documents are not indexed. This usually results from unsupported file formats or file content parsing timeouts, preventing correct document chunking and vectorization.
- Using a general embedding model leads to abnormally high or low similarity scores, resulting in inaccurate search results. This stems from the general model's insufficient semantic understanding of specialized medical vocabulary, failing to distinguish subtle semantic differences.
- Query results fail to effectively filter or sort data by specific time periods or dosage ranges. This may occur because these structured metadata were not specially processed as queryable or sortable attributes during vectorization or indexing.
How to Verify Correct Configuration
- Query a set of drug reports containing known adverse reactions. Check if the recalled results include all key adverse event descriptions, drug information, and patient characteristics.
- Use different specialized medical terms as query words. Verify if the recalled results accurately match relevant documents. Observe the similarity score distribution to ensure expected discriminative power.
- Upload a new drug label or adverse event report. Check if it is indexed promptly. Verify through queries that the new data is retrievable, confirming the incremental update mechanism functions correctly.
The values provided are common starting points. Measure them against specific samples to ensure optimal performance for individual use cases.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.