Data Characteristics in this Category
Real-World Evidence (RWE) products source data from various origins, including Electronic Health Records (EHR), medical claims data, patient registries, wearable device data, genomic data, and social media information. Data update frequencies vary; some EHR data may update in real-time, while large study cohort data might update quarterly or annually in batches. Document structures are complex, containing unstructured clinical notes, semi-structured lab reports, and structured diagnostic codes. Fields and units are highly specialized, such as International Classification of Diseases codes (ICD-10), Medical Subject Headings (MeSH), generic drug names and dosage units (mg/kg), and various biomarker test values. Data volume is typically large, with significant redundancy, missing values, and inconsistencies.
Constraints Imposed by These Characteristics on Vector Models and Indexing
The high heterogeneity of RWE data necessitates complex data cleaning and standardization before vector model construction. Unstructured text (e.g., clinical notes) requires advanced natural language processing techniques for entity recognition and relation extraction to derive meaningful features. Inconsistent data update frequencies demand indexing strategies that support incremental updates, avoiding resource consumption from full rebuilds. The abundance of specialized terminology and abbreviations challenges general-purpose vector models, potentially requiring domain-specific pre-training or fine-tuning. Furthermore, common sensitive information in RWE data, such as patient privacy, requires anonymization or encryption mechanisms during indexing. The specificity of fields and units requires vector models to distinguish semantically similar but contextually different entities, such as the same drug administered via different routes.
Configuration Guidelines
| Configuration Item | Suggested Value | Rationale |
|---|---|---|
Chunk size (Chunk Length) | 500–800 characters | Balances semantic completeness with vector model processing efficiency |
Overlap Length | 50–100 characters | Ensures contextual continuity, reduces semantic breaks |
embedding_model | Domain-specific or fine-tuned model | RWE data is highly specialized; general models lack sufficient semantic understanding |
Similarity threshold (Similarity Threshold) | Calibrated by measurement 0.75–0.85 | Balances recall and precision, prevents irrelevant information interference |
Recall count (Recall Count) | Top 10–20 items | Balances retrieval efficiency with result coverage, reduces unnecessary computation |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | RWE documents are often large, requiring longer processing times |
Three Common Pitfalls
- Knowledge base status shows "Indexing" for an extended period: This typically results from file parsing timeouts. RWE documents often contain complex formats and embedded objects, leading to prolonged parsing.
- Search results show abnormally high or low semantic similarity: This can occur if the
Similarity threshold(Similarity Threshold) is not adjusted to the new model's output range after switching vector models. For example, some models output similarity values far exceeding0-1. - Retrieval results contain a large amount of irrelevant information: This is common when
Recall count(Recall Count) is set too high, or theembedding_modelfails to effectively capture RWE domain-specific semantic relationships.
How to Confirm Proper Configuration
- Select a representative set of RWE question-answer pairs. Evaluate the accuracy and relevance of retrieval results through manual assessment or automated testing. Adjust the
Similarity threshold(Similarity Threshold) based on evaluation outcomes. - Monitor knowledge base indexing logs. Ensure no
PARSE_FILE_TIMEOUT_SECONDS-related errors or warnings appear, confirming all documents are successfully indexed. - Perform retrieval tests on different types of RWE data (e.g., clinical notes, genetic reports). Verify that the
embedding_modelworks effectively across all data types and that recalled content covers key information. - Periodically sample indexed data. Verify that fields and units are correctly identified and embedded, for example, by checking vector representations of specific medical terms or drug names.
The values provided are common starting points. Measure them against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.