Data Characteristics in this Category
Real-world evidence (RWE) data in pharmacovigilance originates from various sources, including Electronic Health Records (EHRs), medical claims databases, patient registries, wearable devices, and social media. Data update frequencies range from daily (for some EHR systems) to quarterly or annually (for large-scale registry studies). Document structures are highly heterogeneous, often containing unstructured free text (e.g., clinician notes, patient self-reports), semi-structured medical terminology codes (e.g., ICD-10, SNOMED CT), and structured data like laboratory results and medication records. Fields and units are highly specialized, such as dosage units (mg, μg/kg), frequencies (QD, BID), durations (days, weeks, months), adverse event terms (MedDRA codes), and complication descriptions. There is also extensive use of abbreviations and domain-specific expressions.
Constraints Imposed by These Characteristics on "Deployment and Upgrade"
The heterogeneity and multi-source nature of RWE data present challenges for data preprocessing and knowledge base construction. Unstructured text requires robust Natural Language Processing (NLP) capabilities for information extraction and standardization, directly impacting computational resource allocation. Frequent data updates necessitate deployment solutions that support incremental updates and version management to ensure knowledge base timeliness and traceability. Medical terminology codes and domain-specific expressions require customized vocabularies and entity recognition models, increasing the complexity of model training and maintenance. Furthermore, the large volume of data makes index building and query efficiency critical bottlenecks, requiring optimization of database and vector store performance during deployment. Maintaining standardization and mapping rules for heterogeneous data fields also demands flexible configuration and upgrade capabilities from the system.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
UPLOAD_FILE_MAX_SIZE | 200 MB | Accommodates the upload of large electronic medical records or research reports, preventing upload failures due to excessive file size. |
maxContext | 1000–1500 characters | Balances contextual understanding of lengthy medical texts with computational cost, ensuring critical information is not truncated. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Addresses situations where complex PDFs or multi-page documents require longer parsing times, preventing parsing failures due to timeouts. |
Chunk size | 400 characters | Optimizes segmentation of unstructured clinical notes and case reports, improving retrieval granularity. |
Recall count | Top 8 entries | Increases the coverage of potentially relevant adverse event information recalled from vast amounts of real-world data. |
Similarity threshold | 0.78 | Precisely matches medical terms and adverse reaction descriptions that have high specialization and subtle differences. |
Three Common Pitfalls
- The container remains in a
startingstate for an extended period after launch, with no clear error messages in the logs. This can be caused by incorrect database connection configurations or a database service that has not started correctly, preventing application initialization. - After uploading a large research report file, the system reports
file parsing failed, but small files parse normally. This is typically due to insufficientPARSE_FILE_TIMEOUT_SECONDSor memory limits, preventing the system from completing complex document parsing within the allotted time. - The recall rate for specific medical terms or disease names in knowledge base queries is unusually low. This can occur if customized vocabularies are not loaded correctly, or if synonyms and variations of medical terms were not adequately considered during index construction.
How to Verify Correct Configuration
- Upload a typical real-world research report containing various data types (e.g., structured tables, unstructured text, medical image links) and verify successful parsing and indexing.
- Execute queries for specific drug adverse reactions. Cross-reference the returned results to ensure they include relevant information from different data sources (e.g., EHR, patient reports) and check the accuracy of information extraction.
- After updating a batch of new real-world data, verify that the knowledge base's incremental update mechanism functions correctly and confirm that new data is included in the retrieval scope through queries.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.