Data Characteristics in this Domain
Pharmacovigilance data originates primarily from post-market adverse event reports, clinical trial data, literature reviews, regulatory documents, and drug labels. This data updates frequently. Adverse event reports may update in real-time or daily, while regulatory documents revise periodically based on requirements. Document structures vary, including structured case report forms, semi-structured free-text reports, and unstructured research papers. Fields cover patient demographics, drug usage, adverse event descriptions, medical terminology codes (e.g., MedDRA or WHODrug), and causality assessments. Fields like dosage, frequency, and duration often include specific units.
Constraints Imposed on Vector Models and Indexing
Pharmacovigilance data diversity places specific demands on vector model selection and indexing strategies. Real-time or high-frequency data sources require support for incremental indexing and fast queries to ensure information timeliness. Complex document structures necessitate more refined text chunking strategies. For example, extract key fields from structured sections and semantically chunk free text. Medical terminology codes require vector models to effectively process specialized vocabulary and abbreviations, and understand their hierarchical relationships. Furthermore, extracting critical information like causality assessments demands high capability from the model to capture contextual semantics. Standardized handling of measurement units, to avoid ambiguity during vectorization, is also an important consideration during index construction.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size | 500–800 characters | Balances semantic completeness with vector model processing capability. Avoids overly long text diluting key information or overly short text losing context. |
Recall count | Top 10 entries | Ensures recall while reducing computational overhead for subsequent re-ranking and processing. Covers the common information density of adverse event reports. |
Similarity threshold | 0.75–0.85 | Balances recall accuracy and false positive rate. Too low may introduce irrelevant results; too high may miss important information. Adjust the specific value based on actual data. |
Rerank result count | Top 3 entries | Focuses on the most relevant results, increasing the density of effective information presented to the user and reducing manual screening costs. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Accommodates parsing time for large or complex documents (e.g., multi-page clinical trial reports), preventing indexing failures due to timeouts. |
vector_model_id | text-embedding-v3-large | Selects a vector model with high semantic understanding and multilingual support, suitable for the diverse text data in pharmacovigilance. |
Three Common Pitfalls
- The knowledge base status remains "indexing" for an extended period. This may occur if the
PARSE_FILE_TIMEOUT_SECONDSparameter is set too low, causing the system to time out when processing large PDFs or complex structured documents. - Search results show excessively high semantic similarity (e.g.,
1.0) and cannot be effectively filtered. This usually results from an inappropriate vector model selection or insufficient text preprocessing, leading to overly similar vectors for different document segments. - Requests for a custom channel's vector model are routed to the LLM. This can happen when the
vector_model_idconfiguration does not match the actual model service, preventing the system from correctly identifying the vector model.
How to Verify Configuration
- Upload representative pharmacovigilance documents. Check if the knowledge base status eventually changes to "completed" and verify the parsed content's completeness.
- Select query statements with known answers. Perform a knowledge base search. Check if the returned
Similarity thresholdandRecall countmeet expectations. Adjust thresholds based on expert judgment. - Use different types of queries (e.g., statements containing medical terms, drug names, or adverse event descriptions). Verify the vector model's understanding of specialized vocabulary. Check if the results in
Rerank result countare the most relevant.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.