Data Characteristics in this Category
Pharmacovigilance data originates from clinical trial reports, real-world evidence (RWE), adverse event reports (ADRs), and drug labels. This data updates frequently. Safety data, especially, continuously generates during ongoing clinical trials. Document structures typically include both structured and unstructured information. Examples include patient demographics, medication history, disease diagnoses, adverse reaction descriptions, and laboratory test results. Unstructured text, such as adverse event narratives, often contains many medical terms, abbreviations, and colloquialisms. Fields involved include drug name, batch number, dosage, administration route, event time, severity, and outcome. Units cover dosage units (mg, g, IU), time units (days, hours), and frequency. Compatibility issues with multiple languages and different measurement systems require handling.
Constraints from these Characteristics on "Vector Models and Indexing"
High update frequency requires vector indexes to support efficient incremental updates. This ensures the timeliness of retrieval results. Mixed structured and unstructured data requires vector models to effectively process different information types. Semantic understanding of unstructured text is crucial. This requires capturing complex medical concepts and relationships. The vast number of medical terms and abbreviations demands higher quality word embeddings. General models may struggle to accurately understand their meaning in a pharmacovigilance context. This can lead to insufficient relevant recall. The presence of multiple languages and different units increases data preprocessing complexity. Unification or effective mapping in the vector space is necessary. Furthermore, accurate identification of key fields like adverse event severity and outcome directly impacts pre-screening accuracy. This requires vector models to distinguish subtle semantic differences.
Configuration Settings
| Configuration Item | Recommended Value | Rationale for Recommendation |
|---|---|---|
Chunk size (Segment Length) | 500–800 characters | Balances contextual completeness with vector model processing efficiency. Avoids excessively long segments that dilute key information. |
Recall count (Recall Count) | Top 10–20 items | Ensures enough candidate results enter the re-ranking stage. Covers potential relevant information and avoids omissions. |
Similarity threshold (Similarity Threshold) | Calibrate based on actual measurements | Balances recall and precision according to specific business scenarios. Tune using a small dataset. |
Rerank result count (Re-ranking Return Count) | Top 3–5 items | Focuses on the most relevant results. Reduces manual review burden and improves pre-screening efficiency. |
embeddingModel | text-embedding-ada-002 or domain-fine-tuned model | Balances general semantic understanding with the specialized vocabulary of the medical domain. Improves vector representation quality. |
PARSE_FILE_TIMEOUT_SECONDS | 300 seconds | Accommodates parsing time for large clinical trial reports or adverse event summary documents. Prevents indexing failures due to timeouts. |
Three Common Mistakes
- Query results show generally low relevance after index construction. This may be because the chosen vector model is not optimized for specialized terminology in the biomedical field, leading to semantic understanding deviations.
- Some adverse event report documents remain in an "indexing" state for a long time, failing to complete indexing. This usually occurs due to document parsing timeouts, especially for PDF files containing many images or complex tables.
- When querying for adverse reactions to a specific drug, the number of recalled items is far less than expected. This may be because the text segmentation granularity is too large, causing key information to be diluted within long, irrelevant text segments.
How to Confirm Proper Configuration
- Select a test set containing known adverse events. Perform retrievals for key symptoms or drug names. Check if the expected documents are included in the recalled results. Adjust
Recall count(Recall Count) andSimilarity threshold(Similarity Threshold) based on business requirements. - Monitor the indexing task queue. Ensure large documents (e.g., PDFs over 10MB) can complete parsing and indexing normally. Check if the
PARSE_FILE_TIMEOUT_SECONDSsetting is sufficient. - Use medical texts of varying lengths and complexities for segment preview. Evaluate the completeness and independence of key information under the
Chunk size(Segment Length) parameter. Ensure each segment contains meaningful semantic units.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.