Data Characteristics
Pharmacovigilance data in medical affairs originates from clinical trial reports, real-world evidence (RWE) data, case reports, regulatory guidelines, and academic literature. This data updates frequently, especially after new drug launches or the emergence of new adverse events. Document structures are diverse, including structured Case Report Forms (CRFs), semi-structured medical journal articles, and unstructured free-text descriptions. Data fields include patient demographics, medication history, disease diagnoses, adverse event occurrence times, severity, outcomes, and physician assessments. Units typically involve dosage (mg, g), frequency (once daily, twice weekly), time (days, weeks, months), and laboratory indicators (mmol/L, U/L).
Constraints Imposed by These Characteristics on Vector Models and Indexing
The high update frequency of pharmacovigilance data requires vector indexes to support efficient incremental updates, ensuring the timeliness of retrieval results. Diverse document structures necessitate processing various text formats, particularly accurately extracting key information from unstructured text. Medical terminology, abbreviations, and synonyms in free text demand advanced semantic understanding from vector models; models must recognize and differentiate subtle medical concepts. Key information like adverse event severity and outcomes is crucial for retrieval accuracy and relevance, requiring effective encoding during vectorization. The need for precise recall results makes a re-ranking step after vector retrieval indispensable, requiring optimization based on business logic.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale for Recommendation |
|---|---|---|
Chunk size (Segment Length) | 500–800 characters (characters) | Ensures each segment contains sufficient context while avoiding information overload that could impact vector quality. |
Chunk Overlap Length (Segment Overlap Length) | 100–150 characters (characters) | Maintains semantic continuity between segments, preventing critical information from being truncated. |
embedding_model | text-embedding-v1 or bge-large-zh-v1.5 | Balances understanding of Chinese medical terminology with model performance. |
Recall count (Recall Count) | 20–30 entries (items) | Provides a sufficient candidate set for subsequent re-ranking, covering potentially relevant information. |
Similarity threshold (Similarity Threshold) | 0.75–0.85 (calibrated by actual measurement) | Balances recall and precision, filtering out irrelevant document snippets. |
Rerank result count (Re-ranked Return Count) | 5–8 entries (items) | Delivers refined and highly relevant results, reducing user reading burden. |
Common Pitfalls
- During index updates, failing to perform incremental updates or experiencing excessively long full re-indexing times, leading to retrieval results that do not reflect the latest data. This occurs due to not enabling incremental indexing or improper incremental strategy configuration.
- Multimodal Embedding model testing fails with error
{"error":{"code":"Invalid, typically due to incorrect API Key configuration or a model name that does not match the actual available service. - In retrieval results, critical information about the severity or outcome of adverse drug events is missing, impacting decision-making. This may be because these structured fields were not effectively encoded during vectorization, or the segmentation strategy caused critical information to be split.
How to Confirm Proper Configuration
- Select a batch of documents containing recently published adverse event information. Verify that retrieval results include this latest information.
- Use queries with clear descriptions of adverse event symptoms. Check if the returned results accurately mention relevant symptoms, drugs, and severity, comparing them against manually judged expected outcomes.
- Adjust the
Similarity threshold(Similarity Threshold) and observe changes in recall count and relevance until high recall is maintained while reducing the appearance of irrelevant information.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.