Data Characteristics
Pharmacovigilance market access data primarily comes from regulatory documents, guidelines, post-market safety reports, drug label change records, and adverse event (AE) or serious adverse event (SAE) databases published by drug regulatory agencies. This data updates frequently, especially when new drugs launch, indications expand, or safety signals emerge. Document formats vary, including structured database records (e.g., MedDRA codes), semi-structured regulatory announcements, and unstructured clinical study reports and literature reviews. Fields and units are highly specialized, such as drug generic names, brand names, ATC classification codes, adverse reaction event names, incidence rates, severity, report dates, and causality assessments. Some fields also involve dosage, administration, patient characteristics (e.g., age, gender, concomitant medications), with units strictly adhering to international standards.
Constraints on Vector Models and Indexing
High-frequency updates of regulatory documents and safety reports require vector indexes to support efficient incremental updates, ensuring information timeliness. Diverse document structures, especially the presence of unstructured text, increase text preprocessing complexity. This necessitates more refined text segmentation strategies to capture key information. Specialized fields and units, along with their complex relationships, mean simple keyword matching is insufficient for recall. Vector models must understand the semantics and contextual relationships of professional terminology. For example, the model needs to recognize synonyms like "myocardial infarction" and "AMI," or semantically understand complex concepts such as "dose-dependent adverse reactions." Furthermore, accurate extraction of key judgments like causality and severity places higher demands on the vector model's semantic understanding capabilities, impacting the accuracy of relevant information recall.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Segment Length) | 500-800 characters (characters) | Balances semantic completeness and vector model processing efficiency. Avoids dilution of key information in long texts and loss of context in short texts. |
Chunk overlap (Segment Overlap) | 50-100 characters (characters) | Ensures key information correlation across paragraphs, improving recall rate. |
Recall count (Recall Count) | 10-20 entries (items) | For complex queries in specialized fields, increases recall scope to capture more potentially relevant document snippets. |
Similarity threshold (Similarity Threshold) | 0.75-0.85 | A higher threshold filters out irrelevant results while retaining matches for semantically similar professional terms. |
Rerank result count (Rerank Return Count) | 3-5 entries (items) | After initial vector screening, further refines sorting to focus on the most relevant key information. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds (seconds) | Addresses the parsing needs for large regulatory documents or clinical reports, ensuring complete file processing. |
Common Pitfalls
- Symptom: Knowledge base indexing progress stalls for an extended period, or indicates that some data is not fully indexed. Reason: File parsing timeout, or file content is too large, leading to a bottleneck in the vector model's processing, especially when handling large PDF regulatory documents.
- Symptom: When querying adverse reactions for a specific drug, results include a large amount of irrelevant general medical information. Reason: The vector model did not sufficiently understand the specialized terms and context in the query, leading to generalized recall, or the segmentation strategy was too coarse.
- Symptom: After updating drug labels or regulatory announcements, relevant query results do not reflect the latest information promptly. Reason: The incremental indexing mechanism did not trigger effectively, or the index update frequency did not match the data source's update rhythm.
How to Verify Configuration
- Select recently updated regulatory documents or safety reports. Perform relevant queries and check if the returned results include key changes from the documents. Verify information timeliness.
- For specific causality descriptions or severity judgments in drug adverse event reports, perform precise queries. Verify if the recall results accurately point to the relevant original text snippets.
- Submit specialized queries containing synonyms or hypernyms, such as querying both "myocardial infarction" and "AMI." Check if the recall results cover relevant information under both expressions and evaluate the impact of the similarity threshold on the results.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.