Data Characteristics in This Category
Real-World Evidence (RWE) quality documents draw data from diverse sources. These include Electronic Health Records (EHR), medical insurance claims data, patient registry systems, mobile health device data, and patient-reported outcomes (PROs). Data typically appears as unstructured text, semi-structured tables, and structured database records. Update frequencies vary from daily (e.g., EHR) to quarterly or annually (e.g., some registry studies). Document structures are complex, encompassing study protocols, ethics approvals, data management plans, statistical analysis plans, and study reports. These documents often contain extensive medical terminology, abbreviations, dosage units (e.g., mg/kg, IU), time units (e.g., weeks, months, years), and statistical symbols (e.g., P-values, CI). Field names may include clinical trial-specific codes such as ICD-10, LOINC, and SNOMED CT.
Constraints from These Characteristics on Reference Source and Traceability
The diverse sources and complex structure of RWE documents challenge reference traceability. Extracting key information from unstructured text, such as specific study conclusions or patient characteristic descriptions, requires precision and linkage to the original source. High-frequency data sources demand timely knowledge base synchronization to ensure reference currency. Unique medical terminology and units in documents require specialized dictionary mapping or contextual understanding during retrieval and referencing to avoid ambiguity. For example, minor differences in dosage units can lead to citation errors. Furthermore, the hierarchical relationships within documents like study protocols and ethics approvals necessitate references that point to specific paragraphs and trace back to the overall document structure. Referencing statistical results requires providing the data source, analysis method, and results to ensure verifiability.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Chunk Length) | 800–1200 characters | Accommodates long RWE document paragraphs, balancing contextual completeness and retrieval granularity. |
Overlap Length | 100–200 characters | Ensures continuity of information across chunks, capturing key connection points. |
Recall count (Recall Count) | Top 5–8 items | Improves recall for complex queries, covering multi-dimensional information points. |
Similarity threshold (Similarity Threshold) | Calibrate by measurement | Balances precision and recall based on the semantic similarity distribution of the specific dataset. |
Rerank result count (Rerank Return Count) | Top 3 items | Focuses on the most relevant information, reduces redundancy, and improves final citation quality. |
MAX_TOKENS_PER_CHUNK | 4000 characters | Adapts to long sentences and complex expressions in medical texts, preventing truncation of critical information. |
Three Common Mistakes
- Symptom: AI-generated content shows red or empty reference numbers. Reason: The knowledge base chunking strategy is inadequate, leading to key information being split or insufficient context, preventing the model from effectively linking to the original text.
- Symptom: In a workflow for RWE-specific questions, only the first knowledge base retrieval node successfully provides references; subsequent nodes fail. Reason: Global variables or context passing are misconfigured, preventing subsequent retrieval nodes from correctly obtaining
datasetIdor query parameters. - Symptom: AI directly outputs knowledge base text with errors in statistical data or drug dosage units. Reason: The knowledge base preprocessing stage did not standardize or entity-recognize RWE-specific medical terms and units. This causes the model to copy directly without contextual understanding during referencing or to introduce deviations during formatted output.
How to Confirm Correct Configuration
- Select an RWE document with clear statistical results, dosage information, or patient characteristic descriptions. Ask questions and check if the AI-generated content accurately references specific paragraphs in the original text. Verify that cited statistical data and dosage units match the original.
- Test a workflow containing multiple knowledge base retrieval nodes. Observe if each node correctly triggers knowledge base retrieval and provides corresponding reference sources. Ensure that critical parameters like
datasetIdare correctly passed within the workflow. - For RWE documents with numerous medical terms and abbreviations, test if the AI can accurately identify and cite relevant definitions or explanations. Check if the cited context includes sufficient medical background information to ensure readability and accuracy of references.
The values provided are common starting points. Measure them against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.