Data Characteristics in This Category
Real-World Study (RWS) data for clinical trial pre-screening primarily originates from Electronic Health Records (EHR), medical insurance claims databases, disease registries, and Patient-Reported Outcomes (PROs). This data typically exists in various forms: unstructured text, semi-structured tables, and structured numerical values. Update frequencies vary; EHR data might update in real-time or daily, while medical insurance claims data might update in monthly or quarterly batches. Document structures are complex. For example, EHR progress notes contain extensive free text, while claims data has fixed coding fields. Fields and units are diverse, including diagnostic codes (e.g., ICD-10), drug codes (e.g., ATC), lab results (e.g., mmol/L, ng/mL), and patient demographic information.
Constraints Imposed by These Characteristics on "Citation and Provenance"
The highly heterogeneous nature and irregular update frequency of RWS data challenge the accuracy and timeliness of citation sources. Information in free text requires precise extraction to avoid miscitations or missing critical context. Integrating and correlating multi-source data demands that the system identify and trace back to the specific origin of the raw data, such as which hospital's medical record or which timestamped insurance record. Data sensitivity (patient privacy) means strict adherence to data governance rules for anonymization during citation and provenance. Furthermore, numerical data like lab results require clear specification of units and normal ranges when cited to avoid ambiguity.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Segment Length) | 500–800 characters | RWS free text is information-dense. Too long dilutes key information, too short loses context. This range helps maintain semantic completeness. |
Recall count (Recall Count) | Top 10–15 entries | Given the complexity of RWS data sources, increasing the recall count improves coverage and compensates for potential fragmentation in single pieces of information. |
Similarity threshold (Similarity Threshold) | 0.75–0.85 | For the professional and rigorous nature of medical text, a higher threshold helps filter for more relevant and accurate citation segments. |
Rerank result count (Rerank Return Count) | Top 5 entries | After reranking, selecting a small number of the most relevant entries reduces redundancy and improves traceability efficiency. |
maxContext | 3000–4000 tokens | In RWS pre-screening scenarios, a sufficiently long context is needed to understand complex medical descriptions and multi-source information associations. This range can accommodate more details. |
Citation Source Display Format | File Name + Page Number/Paragraph Number | Clearly indicates the specific location in the original document from which the cited content is drawn, facilitating quick location and verification by engineers. |
Three Common Mistakes
- Too many documents cited in the answer, leading to excessively long generated answers and exceeding token limits. This occurs when
Recall count(Recall Count) orRerank result count(Rerank Return Count) are set too high, failing to effectively filter non-core information. - After publishing a knowledge base page, cited literature does not display specific sources. This is typically because the
Citation Source Display Formatfeature is not enabled or incorrectly configured in the knowledge base settings. - When dealing with RWS data containing many English medical terms, the RAG model generates citation prompts in Chinese. This is because the
Citation Prompt Templatehas not been adjusted for a multilingual environment or lacks corresponding language switching logic.
How to Confirm Correct Configuration
- Select multiple typical clinical pre-screening cases. Verify that the generated answers accurately cite key information from RWS data and check that the cited
File NameandPage Number/Paragraph Numberare clear and traceable. - Simulate RWS data update scenarios. Observe the knowledge base indexing update mechanism and verify if new data is cited promptly, while also checking if old data citations become invalid.
- In a test environment, deliberately introduce RWS data containing sensitive information. Verify that the system adheres to anonymization principles during citation and provenance. Check if
Citation Sourceleaks patient privacy. - Through API or UI, inspect the return structure of citation sources. Ensure that the
Citation Source Display Formataligns with expectations, for example, including file ID, version number, and other traceable metadata.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.