Data Characteristics in This Category
Real-World Evidence (RWE) in pharmacovigilance draws from diverse data sources. These include Electronic Health Records (EHRs), insurance claims databases, patient registries, and wearable device data. Data typically exists as unstructured text, semi-structured tables, and structured fields. Update frequencies vary: some data, like outpatient records, update in real-time, while large insurance databases update quarterly or annually in batches. Document lengths differ significantly, ranging from case summaries of tens of characters to clinical reports spanning thousands. Fields and units are highly domain-specific, such as drug generic names, batch numbers, dosage units (mg, IU), routes of administration, adverse event terms (using MedDRA coding), complications, and laboratory test results with their units (mmol/L, U/L).
Constraints Imposed by These Characteristics on Knowledge Base Retrieval and Recall
The breadth and heterogeneity of RWE data sources require a knowledge base capable of integrating various data formats effectively. Adverse event descriptions in unstructured text demand precise semantic understanding to prevent recall bias from synonyms, abbreviations, or colloquialisms. Structured fields, such as MedDRA codes or lab indicators, need preprocessing to maintain semantic integrity and avoid incorrect tokenization. Varying data update frequencies necessitate flexible knowledge base synchronization strategies; real-time sources require low-latency updates, while batch-updated data can use periodic full or incremental synchronization. Document length differences challenge chunking strategies; overly long documents may lead to individual chunks containing too much information, while overly short ones might lose context. Furthermore, the presence of domain-specific fields and units requires retrieval models to identify and differentiate these specialized terms, for example, distinguishing "unit" from "dosage unit."
Configuration Guidelines
| Configuration Item | Suggested Value | Rationale for This Value |
|---|---|---|
Chunk Length | 800–1200 characters | RWE documents vary greatly in length; this range balances contextual completeness and retrieval efficiency. |
Overlap Length | 100 characters | Ensures contextual continuity at chunk boundaries, reducing information loss. |
Recall Count | Top 8 | Given the complexity of RWE queries, increasing the recall count appropriately covers more potentially relevant information. |
Similarity Threshold | Calibrate by measurement | Requires adjustment based on specific RWE datasets and query types through testing; typically between 0.7–0.8. |
Rerank Return Count | Top 5 | Further enhances relevance from recall results using a reranking model, focusing on key information. |
File Parse Timeout | 300 seconds | Allows sufficient parsing time for large RWE clinical reports or EHR files. |
Three Common Pitfalls
- Retrieval results contain a large amount of irrelevant data. This occurs when the
Similarity Thresholdis set too low, leading to generalized recall. - Certain critical adverse event information is not recalled, resulting in missing query results. This might stem from a
Chunk Lengththat is too short, truncating context, or an insufficientRecall Count. - API call results differ from online chat interface results, with online chat performing better. This usually happens when the API call does not correctly pass or uses
promptormodel_idparameters that do not match the online chat environment.
How to Verify Correct Configuration
- For typical adverse event queries, check if retrieved documents contain all known relevant information and evaluate their
Relevance Score. - Upload RWE documents containing professional fields like MedDRA codes and laboratory indicators. Perform retrieval to confirm these fields are correctly identified and recalled.
- Compare retrieval results under different
Recall CountandSimilarity Thresholdvalues to assess recall and precision, then select a configuration that achieves balance.
Note: The values provided are common starting points and should be measured against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.