Data Characteristics
Pharmacovigilance regulation data originates from adverse event reports submitted by pharmaceutical manufacturers, medical institutions, patients, and regulatory bodies. It also includes drug labels, risk management plans, post-marketing study reports, pharmacovigilance system documents (e.g., SOPs, guidelines), and relevant laws and regulations. Data update frequencies vary; adverse event reports may be continuous, while SOPs and regulations update less frequently. Document structures are diverse, including structured reports (e.g., XML E2B files), semi-structured text (e.g., SOPs, guidelines), and unstructured text (e.g., patient descriptions). Core fields include drug name, batch number, adverse event description, reporter information, event date, actions taken, and outcome. Units involve dosage (mg, g), frequency (times/day), and time (days, months, years).
Constraints on Vector Models and Indexing
Pharmacovigilance data's diverse and heterogeneous nature requires strict standardization and cleaning before vectorization. Semi-structured and unstructured documents need more refined text chunking strategies to ensure semantic completeness. Adverse event descriptions often contain extensive medical terminology, abbreviations, and colloquialisms, demanding domain-adapted pre-trained models. The long text nature of regulatory documents can lead to individual chunks being too long, diluting core information, or too short, losing context. Inconsistent update frequencies necessitate incremental indexing capabilities to avoid frequent full rebuilds. Sensitive information in reports (e.g., patient identity) requires anonymization before vectorization to prevent leakage. Identifying and normalizing field units is crucial for retrieval accuracy and comparability.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
chunk_size | 500–800 characters | Balances chunk completeness and retrieval efficiency, adapts to various document structures |
overlap_size | 50–100 characters | Ensures semantic continuity between chunks, prevents loss of context |
embedding_model | dengcao/Qwen3-Embedding-8B | Optimized for Chinese medical domain, better understands professional terminology |
similarity_threshold | Calibrate by actual measurement | Based on precision and recall curves to ensure relevance |
recall_top_k | top 10 | Ensures sufficient recall range, covers potentially relevant information |
rerank_model | dengcao/Qwen3-Rerank | Improves ranking quality of recalled documents, filters out low-relevance results |
Common Pitfalls
- Semantic fragmentation after chunking. Retrieval results fail to provide complete answers. This occurs when
chunk_sizeis too small or document paragraph structure is not considered. - Retrieval results contain excessive irrelevant information. This occurs when
similarity_thresholdis set too low, recalling chunks with low relevance to the query. - Index updates take too long. Each new document requires a long wait. This occurs when incremental indexing strategies are not configured or bottlenecks exist in the document processing pipeline.
Configuration Validation
- Select a set of standard pharmacovigilance-related questions. Evaluate the proportion of correct answers included in the recall results.
- Examine the contextual completeness of returned chunks in retrieval results. Determine if critical information is truncated.
- Monitor index service resource utilization. Confirm that index update time after adding new documents is within an acceptable range.
- Adjust
similarity_threshold. Observe changes in the number of retrieval results. Combine with manual judgment to determine an appropriate threshold range.
The values provided are common starting points and should be measured against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.