Real-World Evidence Data Characteristics
Real-World Evidence (RWE) data for registration and declaration typically originates from various sources. These include Electronic Health Records (EHR), medical insurance claims databases, disease registries, patient-reported outcome (PRO) data, and wearable device data. This data often exists as unstructured text, semi-structured tables, and structured database records. Update frequencies vary; EHR and claims data might update daily or weekly, while disease registry data could update quarterly or annually. Document types are diverse, encompassing clinical records, lab reports, medical imaging reports, drug prescriptions, and follow-up records. Field content is complex, including diagnostic codes (e.g., ICD-10), drug codes (e.g., ATC codes), laboratory indicators (with units like mg/dL, mmol/L), patient demographic information, and disease progression descriptions.
Constraints from "Referencing and Tracing" Due to These Characteristics
The diversity and dynamic nature of RWE data impose specific requirements on referencing and tracing. Unstructured text referencing requires precision down to the paragraph or sentence level to support complex clinical descriptions and expert judgments. Multiple heterogeneous data sources mean a single knowledge base might not cover all relevant information, necessitating cross-knowledge base retrieval and referencing. Dynamically updated data requires knowledge bases to have version management capabilities, ensuring reference timeliness. Discrepancies in field and unit standardization require additional semantic matching and unit conversion during referencing to avoid misinterpretation. Furthermore, data involving patient privacy requires strict access controls to ensure compliance during the referencing process. Document length and complexity also affect chunking strategies and recall efficiency; excessively long chunks reduce reference precision, while overly short chunks might lose context.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk Length | 800–1200 characters | Balances context completeness and retrieval efficiency, adapting to the paragraph structure of clinical documents. |
Chunk Overlap Length | 100–150 characters | Ensures continuity of information across chunks, preventing critical information from being truncated. |
Recall Count | Top 10–15 entries | Considers the richness of RWE data sources, increasing recall to cover more potentially relevant information. |
Similarity Threshold | 0.75–0.85 | Ensures the precision of recalled content, filtering out irrelevant background information. |
maxContext | 4000–6000 tokens | Accommodates the complexity of RWE reports, providing sufficient context for the model to understand and generate. |
Reference Display Format | Document Name:Page Number:Paragraph Index | Provides a precise reference path, facilitating manual verification and data traceability. |
Common Pitfalls
- Symptom: AI response lacks critical information from the knowledge base, but the reference list shows relevant documents. Cause:
Similarity Thresholdis set too high. Relevant documents are recalled but not sufficiently utilized by the model. - Symptom: API call returns
408 Request Timeoutor504 Gateway Timeouterrors. Cause:maxContextparameter is set too large, causing the model processing time to exceed gateway or proxy timeout limits. - Symptom: AI response references outdated data, inconsistent with the latest version. Cause: The knowledge base lacks an effective data version management mechanism, failing to index the latest updated data.
Verification of Configuration
- Select a test document containing the latest RWE data. Ask a question requiring a specific numerical reference from it. Check if the AI response includes the number and if the reference source
Document Name:Page Number:Paragraph Indexpoints to the correct location. - Choose a test document with complex clinical descriptions. Evaluate the AI's ability to summarize or analyze it. Check if the response includes sufficient references to support its arguments.
- Simulate a data update scenario by uploading a new version of RWE data. Then, ask a question related to the updated content. Confirm that the AI references the latest version of the data.
The values provided are common starting points. Measure against your own samples for optimal configuration.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.