RWE Data Characteristics
Real-World Evidence (RWE) policy data originates from various sources, including electronic health records (EHRs) from healthcare institutions, insurance claims databases, disease registries, and patient-reported outcomes (PRO) data. This data typically exists in unstructured text, semi-structured tables, and structured numerical formats. Update frequencies vary; EHR data might update in real-time, while insurance claims data usually imports in quarterly or annual batches. Document structures are complex, containing medical terminology, abbreviations, and clinical pathway descriptions. Fields and units are diverse; for example, laboratory test results may involve multiple units like mmol/L or ng/mL, and diagnostic information is primarily based on International Classification of Diseases (ICD) codes or free-text descriptions.
Constraints Imposed by RWE Data on Deployment and Upgrade
The diverse sources and varying update frequencies of RWE data necessitate a highly flexible data ingestion module during deployment. This module must adapt to multiple data source interfaces and support incremental update mechanisms. Unstructured text and complex document structures require robust text parsing and entity recognition capabilities to accurately extract critical information such as disease diagnoses, treatment plans, and adverse events. Diverse fields and units demand that the question-answering system perform unit conversion and semantic alignment during knowledge graph construction and answer generation to prevent misunderstandings caused by inconsistent units. During system upgrades, particular attention must be paid to data schema change compatibility, ensuring that new versions can correctly process existing data and smoothly transition to new data structures.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
maxContext | 4000 | RWE documents are often lengthy, requiring a larger context window to capture complete information. |
Segment Length | 800–1200 characters | Balances semantic completeness with segment retrieval efficiency, avoiding overly long or short text segments. |
Similarity Threshold | 0.78–0.85 | Balances recall precision and recall rate, reducing irrelevant results while ensuring critical policy terms are retrieved. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | File parsing can be time-consuming when processing large or structurally complex PDF documents (e.g., clinical study reports). |
Retrieval Count | Top 10 | Given the precision requirements for policy Q&A, retrieve more relevant items for the reranking module to filter. |
Reranked Return Count | Top 3 | Reduces the model's processing burden while ensuring the most relevant policy terms are returned as answer basis. |
Three Common Mistakes
- File parsing timeout, with logs showing
ERROR: File parsing timed out. This may be due toPARSE_FILE_TIMEOUT_SECONDSbeing set too low, preventing the processing of large or structurally complex PDF documents. - Question-answering results lack critical dosage or unit information. This often occurs when the knowledge base construction fails to effectively perform entity recognition and association for numerical values and units in the text, leading to information loss.
- After system startup, the interface displays a
Service Unavailableerror. This may be due to incorrectdepends_onconfiguration in the Docker Compose file, leading to an incorrect service startup order where the database or model services are not ready.
Verification Steps
- Upload an RWE report containing complex tables and medical terminology. Check if its segmentation is reasonable and if critical information is correctly extracted and indexed.
- Ask questions about drug dosages and adverse event handling processes for a specific disease treatment policy. Verify the accuracy and completeness of the system's returned answers.
- Simulate data source updates, such as importing a new batch of insurance claims data. Observe if the incremental update of the knowledge base completes as expected and if updated Q&A results reflect the latest information.
The values provided are common starting points and should be measured against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.