Real-World Evidence Pharmacovigilance: Citation and Traceability

Real-World Evidence (RWE) in pharmacovigilance uses diverse data sources. These include Electronic Health Records (EHR), insurance claims databases

Data Characteristics

Real-World Evidence (RWE) in pharmacovigilance uses diverse data sources. These include Electronic Health Records (EHR), insurance claims databases, spontaneous adverse event reporting systems, patient registries, mobile health app data, and wearable device data.

Data formats vary. Non-structured text appears in clinical notes or physician notes. Semi-structured tables contain diagnostic codes or medication information from claims records. Structured numerical data includes lab results or vital signs.

Data update frequencies differ. EHRs and spontaneous reporting systems may update daily or in real-time. Large claims databases typically update monthly or quarterly.

Document structures are complex. A complete patient record might include admission notes, ward rounds, lab reports, imaging reports, and discharge summaries. Each section has a specific format and information organization.

Fields and units are highly standardized for medical terminology, such as ICD-10 diagnostic codes and ATC drug classification codes. However, free-text descriptions contain many non-standard expressions, abbreviations, and colloquialisms.

Constraints for Citation and Traceability

Diverse and unstructured RWE data presents several challenges for citation and traceability.

First, data heterogeneity requires the knowledge base to handle multiple modalities. It must effectively parse text, tables, and structured numerical data, then index them uniformly.

Second, inconsistent data update frequencies demand real-time capabilities from the knowledge base. This prevents the citation of outdated information.

Large data volumes and complex document structures increase the difficulty of knowledge segmentation and vectorization. Fine-grained control over segmentation is necessary to ensure recall accuracy.

Non-standard expressions in free text reduce the accuracy of keyword matching and semantic similarity recall. This requires stronger semantic understanding or preprocessing mechanisms.

Finally, patient privacy is critical. Data anonymization and compliance requirements are extremely high. The system must not leak sensitive information when displaying citations. It must also trace back to the original anonymized data. The path to the original data source must be clear for engineers to troubleshoot and verify data.

Configuration Settings

Configuration ItemRecommended ValueRationale
Chunk size500–800 charactersBalances common RWE report paragraph lengths with semantic integrity. Avoids redundancy from overly long segments and loss of context from overly short segments.
Recall countTop 8–12 entriesRWE data is highly correlated. Recalling more contextual information covers potential adverse event associations.
Similarity threshold0.75–0.85This range balances precise recall with generalization capability, considering subtle differences in medical terminology and the ambiguity of free text.
Rerank result countTop 3 entriesAfter re-ranking model optimization, the top few results usually contain the most relevant and high-quality citations, reducing user reading burden.
PARSE_FILE_TIMEOUT_SECONDS600 secondsRWE report files, especially EHR exports, can contain large amounts of text, requiring longer parsing times.
MAX_EMBEDDING_BATCH_SIZECalibrate by actual measurementEnsures that memory overflow or inefficient batch processing is avoided when vectorizing large-scale medical texts.

Common Pitfalls

  • The returned answer does not match the cited source content. System logs show knowledge_base_recall_empty or similarity_score_low. This occurs due to improper knowledge base segmentation granularity or a similarity threshold set too high, preventing relevant knowledge points from being recalled.
  • Citations are unclear. Only the document name is provided, making it impossible to trace the specific location or paragraph in the original text. Users cannot verify the information. This happens when the knowledge base does not save enough contextual metadata during indexing, or the frontend display logic does not fully utilize this metadata.
  • File upload or parsing times out when processing large RWE reports, displaying upload_file_timeout or parsing_error. This indicates that system parameters like PARSE_FILE_TIMEOUT_SECONDS are set too low for the complexity and size of RWE reports.

Validation Steps

  • Select multiple RWE reports containing typical adverse event descriptions. Ask relevant questions. Check if the returned answers accurately cite specific paragraphs from the reports and can trace back to the corresponding location in the original text.
  • Choose RWE data containing medical abbreviations, synonyms, or ambiguous descriptions. After querying, check if the recalled knowledge snippets correctly interpret these non-standard expressions and recall relevant standardized content.
  • Monitor recall_count and similarity_score fields in system logs. Ensure that the number of recalls and similarity scores meet expectations across different query scenarios. Avoid a large number of low-score recalls or empty recalls.
  • Test uploading an RWE report with content close to the UPLOAD_FILE_MAX_SIZE limit. Ensure the file is parsed and indexed successfully without timeout or parsing errors.

The values provided are common starting points and should be measured against the reader's own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.