Referencing and Tracing for Real-World Evidence Registration and Declaration Materials

Real-World Evidence (RWE) data for registration and declaration typically originates from various sources. These include Electronic Health Records

Real-World Evidence Data Characteristics

Real-World Evidence (RWE) data for registration and declaration typically originates from various sources. These include Electronic Health Records (EHR), medical insurance claims databases, disease registries, patient-reported outcome (PRO) data, and wearable device data. This data often exists as unstructured text, semi-structured tables, and structured database records. Update frequencies vary; EHR and claims data might update daily or weekly, while disease registry data could update quarterly or annually. Document types are diverse, encompassing clinical records, lab reports, medical imaging reports, drug prescriptions, and follow-up records. Field content is complex, including diagnostic codes (e.g., ICD-10), drug codes (e.g., ATC codes), laboratory indicators (with units like mg/dL, mmol/L), patient demographic information, and disease progression descriptions.

Constraints from "Referencing and Tracing" Due to These Characteristics

The diversity and dynamic nature of RWE data impose specific requirements on referencing and tracing. Unstructured text referencing requires precision down to the paragraph or sentence level to support complex clinical descriptions and expert judgments. Multiple heterogeneous data sources mean a single knowledge base might not cover all relevant information, necessitating cross-knowledge base retrieval and referencing. Dynamically updated data requires knowledge bases to have version management capabilities, ensuring reference timeliness. Discrepancies in field and unit standardization require additional semantic matching and unit conversion during referencing to avoid misinterpretation. Furthermore, data involving patient privacy requires strict access controls to ensure compliance during the referencing process. Document length and complexity also affect chunking strategies and recall efficiency; excessively long chunks reduce reference precision, while overly short chunks might lose context.

Configuration Settings

Configuration ItemRecommended ValueRationale
Chunk Length800–1200 charactersBalances context completeness and retrieval efficiency, adapting to the paragraph structure of clinical documents.
Chunk Overlap Length100–150 charactersEnsures continuity of information across chunks, preventing critical information from being truncated.
Recall CountTop 10–15 entriesConsiders the richness of RWE data sources, increasing recall to cover more potentially relevant information.
Similarity Threshold0.75–0.85Ensures the precision of recalled content, filtering out irrelevant background information.
maxContext4000–6000 tokensAccommodates the complexity of RWE reports, providing sufficient context for the model to understand and generate.
Reference Display FormatDocument Name:Page Number:Paragraph IndexProvides a precise reference path, facilitating manual verification and data traceability.

Common Pitfalls

  • Symptom: AI response lacks critical information from the knowledge base, but the reference list shows relevant documents. Cause: Similarity Threshold is set too high. Relevant documents are recalled but not sufficiently utilized by the model.
  • Symptom: API call returns 408 Request Timeout or 504 Gateway Timeout errors. Cause: maxContext parameter is set too large, causing the model processing time to exceed gateway or proxy timeout limits.
  • Symptom: AI response references outdated data, inconsistent with the latest version. Cause: The knowledge base lacks an effective data version management mechanism, failing to index the latest updated data.

Verification of Configuration

  • Select a test document containing the latest RWE data. Ask a question requiring a specific numerical reference from it. Check if the AI response includes the number and if the reference source Document Name:Page Number:Paragraph Index points to the correct location.
  • Choose a test document with complex clinical descriptions. Evaluate the AI's ability to summarize or analyze it. Check if the response includes sufficient references to support its arguments.
  • Simulate a data update scenario by uploading a new version of RWE data. Then, ask a question related to the updated content. Confirm that the AI references the latest version of the data.

The values provided are common starting points. Measure against your own samples for optimal configuration.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.