Data Characteristics
Real-World Evidence (RWE) in pharmacovigilance uses diverse data sources. These include Electronic Health Records (EHR), insurance claims databases, spontaneous adverse event reporting systems, patient registries, mobile health app data, and wearable device data.
Data formats vary. Non-structured text appears in clinical notes or physician notes. Semi-structured tables contain diagnostic codes or medication information from claims records. Structured numerical data includes lab results or vital signs.
Data update frequencies differ. EHRs and spontaneous reporting systems may update daily or in real-time. Large claims databases typically update monthly or quarterly.
Document structures are complex. A complete patient record might include admission notes, ward rounds, lab reports, imaging reports, and discharge summaries. Each section has a specific format and information organization.
Fields and units are highly standardized for medical terminology, such as ICD-10 diagnostic codes and ATC drug classification codes. However, free-text descriptions contain many non-standard expressions, abbreviations, and colloquialisms.
Constraints for Citation and Traceability
Diverse and unstructured RWE data presents several challenges for citation and traceability.
First, data heterogeneity requires the knowledge base to handle multiple modalities. It must effectively parse text, tables, and structured numerical data, then index them uniformly.
Second, inconsistent data update frequencies demand real-time capabilities from the knowledge base. This prevents the citation of outdated information.
Large data volumes and complex document structures increase the difficulty of knowledge segmentation and vectorization. Fine-grained control over segmentation is necessary to ensure recall accuracy.
Non-standard expressions in free text reduce the accuracy of keyword matching and semantic similarity recall. This requires stronger semantic understanding or preprocessing mechanisms.
Finally, patient privacy is critical. Data anonymization and compliance requirements are extremely high. The system must not leak sensitive information when displaying citations. It must also trace back to the original anonymized data. The path to the original data source must be clear for engineers to troubleshoot and verify data.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size | 500–800 characters | Balances common RWE report paragraph lengths with semantic integrity. Avoids redundancy from overly long segments and loss of context from overly short segments. |
Recall count | Top 8–12 entries | RWE data is highly correlated. Recalling more contextual information covers potential adverse event associations. |
Similarity threshold | 0.75–0.85 | This range balances precise recall with generalization capability, considering subtle differences in medical terminology and the ambiguity of free text. |
Rerank result count | Top 3 entries | After re-ranking model optimization, the top few results usually contain the most relevant and high-quality citations, reducing user reading burden. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | RWE report files, especially EHR exports, can contain large amounts of text, requiring longer parsing times. |
MAX_EMBEDDING_BATCH_SIZE | Calibrate by actual measurement | Ensures that memory overflow or inefficient batch processing is avoided when vectorizing large-scale medical texts. |
Common Pitfalls
- The returned answer does not match the cited source content. System logs show
knowledge_base_recall_emptyorsimilarity_score_low. This occurs due to improper knowledge base segmentation granularity or a similarity threshold set too high, preventing relevant knowledge points from being recalled. - Citations are unclear. Only the document name is provided, making it impossible to trace the specific location or paragraph in the original text. Users cannot verify the information. This happens when the knowledge base does not save enough contextual metadata during indexing, or the frontend display logic does not fully utilize this metadata.
- File upload or parsing times out when processing large RWE reports, displaying
upload_file_timeoutorparsing_error. This indicates that system parameters likePARSE_FILE_TIMEOUT_SECONDSare set too low for the complexity and size of RWE reports.
Validation Steps
- Select multiple RWE reports containing typical adverse event descriptions. Ask relevant questions. Check if the returned answers accurately cite specific paragraphs from the reports and can trace back to the corresponding location in the original text.
- Choose RWE data containing medical abbreviations, synonyms, or ambiguous descriptions. After querying, check if the recalled knowledge snippets correctly interpret these non-standard expressions and recall relevant standardized content.
- Monitor
recall_countandsimilarity_scorefields in system logs. Ensure that the number of recalls and similarity scores meet expectations across different query scenarios. Avoid a large number of low-score recalls or empty recalls. - Test uploading an RWE report with content close to the
UPLOAD_FILE_MAX_SIZElimit. Ensure the file is parsed and indexed successfully without timeout or parsing errors.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.