Real-World Study (RWS) Data Characteristics
Real-World Study (RWS) data primarily originates from Electronic Health Records (EHR), medical insurance claims databases, disease registries, and patient-reported outcome (PRO) data. This data is typically heterogeneous, containing both structured data (e.g., ICD-10 diagnosis codes, ATC drug codes, laboratory test results) and unstructured text (e.g., outpatient medical records, inpatient records, imaging reports). Data update frequencies vary; some medical insurance data may update monthly or quarterly in batches, while EHR data is generated in real-time. Document structures are complex. For instance, medical record text may contain numerous medical abbreviations, colloquialisms, and custom templates from different hospitals. Regarding fields and units, various units of measurement exist (e.g., mg, g, mmol/L), often lacking unified standardization, which can lead to missing or inconsistent units.
Constraints Imposed by These Characteristics on Workflow Orchestration
The heterogeneity of RWS data requires workflows to have robust data preprocessing capabilities. This includes adapting diverse parsers for different data sources and integrating Natural Language Processing (NLP) modules for entity recognition and relationship extraction to handle unstructured text. Inconsistent data update frequencies necessitate workflows that support scheduled triggers and incremental update strategies to avoid reprocessing historical data. Complex document structures and non-standardized fields challenge text segmentation and vectorization, requiring more refined context management and longer segment lengths to preserve the integrity of medical semantics. The issue of inconsistent units demands the introduction of unit standardization steps during the data cleaning phase and dimensional consistency checks in subsequent analyses to prevent misinterpretation due to unit discrepancies. These constraints dictate that workflow orchestration must be highly flexible and configurable to accommodate the diverse characteristics of RWS data.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
maxContext | 4096 | Ensures capacity for critical medical descriptions and contextual information within RWS reports. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Addresses the complexity and time-consuming parsing of RWS report files (e.g., PDF medical records). |
Chunk size | 800–1200 characters | Preserves the integrity of medical entities and their associated descriptions in RWS reports. |
Similarity threshold | 0.75 | Improves the precision of RWS data recall, reducing interference from irrelevant information. |
Recall count | Top 10 entries | Ensures coverage of multiple relevant data points and arguments potentially involved in RWS reports. |
Rerank result count | Top 5 entries | Further filters the most relevant RWS snippets based on the initial recall. |
Common Pitfalls
- During workflow execution, a large amount of low-quality content unrelated to the query is returned. The symptom is that the
recall_itemslist contains significant noise. This may be due toSimilarity threshold(similarity threshold) being set too low, failing to effectively filter out non-core information. - When processing RWS reports, some critical medical information (e.g., diagnoses or treatment plans) is not correctly extracted, leading to missing important content in the final generated report. This may be due to
Chunk size(segment length) being set too short, truncating medical descriptions with complete semantic meaning. - API calls return a
400 Bad Requesterror, indicating missing required parameters, even if these parameters are not explicitly mentioned in the API documentation. This may occur if, during workflow orchestration, an APP (e.g., a data cleaning APP) still checks for global mandatory variables from the previous APP when transitioning to the next APP.
How to Verify Configuration
- Select typical RWS report samples, run them through the workflow, and inspect the
parsed_chunkscontent to evaluate if segmentation is reasonable and critical information is complete. - For specific queries, execute the workflow and review the
recall_itemslist to confirm that the recalled RWS data snippets are highly relevant to the query intent and to verify their accuracy. - Simulate various abnormal data input scenarios (e.g., RWS data missing critical fields), run the workflow, and check error logs or output status codes to ensure that error handling mechanisms trigger as expected.
The values provided are common starting points and should be measured against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.