Data Characteristics
CSO (Contract Sales Organization) pharmacovigilance data originates from clinical trial reports, real-world study data, post-market surveillance data, and adverse event reporting systems from partner pharmaceutical companies. Data updates frequently, typically daily or weekly. Updates are more intensive during initial drug launch or when new adverse event signals emerge. Document structures are diverse, including standardized ICH E2B format reports, unstructured medical text (e.g., handwritten doctor's notes, patient interview records), semi-structured database exports (e.g., CSV, XML), and PDF literature. Fields include patient demographics, medication history, adverse event descriptions (including MedDRA codes), event time, severity, outcome, and causality assessment. Time fields are precise to the day or hour/minute. Dosage units vary (mg, g, ml, IU, etc.) and often involve complex medical terminology and abbreviations.
Constraints on Knowledge Base Retrieval and Recall
The diversity and high update frequency of CSO pharmacovigilance data challenge the real-time nature and accuracy of the knowledge base. The mix of unstructured and semi-structured data requires the knowledge base to effectively extract and integrate multimodal information. High update frequency means the knowledge base must support incremental updates and rapid index reconstruction to prevent retrieval bias from outdated data. Complex medical terminology, abbreviations, and diverse dosage units require the embedding model to have strong semantic understanding, identifying synonyms, near-synonyms, and hierarchical relationships to reduce recall omissions due to inconsistent terminology. Additionally, precise retrieval requirements for critical fields like time and dosage mean simple keyword matching is insufficient. Entity recognition and relationship extraction techniques are necessary to ensure retrieval accuracy and avoid interference from irrelevant information.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Segment Length) | 500–800 characters | Medical texts have strong contextual relevance. Segments that are too short can lose semantic meaning; segments that are too long add irrelevant information, affecting recall accuracy. |
Chunk Overlap Length (Segment Overlap Length) | 100 characters | Ensures semantic coherence at segment boundaries, preventing critical information from being cut off. |
embedding_model | text-embedding-3-large | Addresses complex medical terminology and multilingual adverse event reports, providing stronger semantic understanding. |
Recall count (Number of Recall Items) | Top 10–15 items | Considering the need for high recall, appropriately increase the number of recalled items, then improve precision through re-ranking. |
Similarity threshold (Similarity Threshold) | Calibrate with actual measurements | Balances recall rate and accuracy based on actual adverse event retrieval scenarios, preventing missed reports and false positives. |
Rerank result count (Number of Re-ranked Items) | 5 items | Provides fine-grained sorting, presenting the most relevant few results to engineers to improve efficiency. |
Common Pitfalls
- Knowledge base query results contain a large amount of irrelevant or duplicate information. This occurs when
Recall count(Number of Recall Items) is set too high orSimilarity threshold(Similarity Threshold) is too low, leading to the recall of many low-relevance segments. - Some adverse event reports cannot be accurately retrieved, even when keywords exist in the original document. This may be because
Chunk size(Segment Length) is too short orChunk Overlap Length(Segment Overlap Length) is insufficient, leading to critical information being cut off or incomplete context. - After changing the
embedding_model, query performance does not improve or even declines. This happens when the already imported knowledge base is notre-indexed, causing it to still use the old model's embedding vectors.
Verification Steps
- Select a batch of representative adverse event queries. Verify that the knowledge base results contain all expected critical information.
- Use different keyword combinations to retrieve the same adverse event report. Observe the consistency and completeness of the recall results.
- After a knowledge base update, compare the same queries before and after the update. Confirm that newly added or modified data can be accurately recalled and assess the timeliness of the recall results.
- Check the accuracy of critical fields in the retrieval results (e.g., drug names, adverse reaction terms, dosage units). Ensure no information bias occurs due to segmentation or embedding model issues.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.