Data Characteristics
Real-World Evidence (RWE) quality documents include research protocols, data management plans, statistical analysis plans, ethics approval documents, informed consent forms, study reports, and publications. These documents are typically in PDF, Word, or structured text formats. They are detailed and highly specialized. Data sources are diverse, including electronic health records, medical claims data, patient registries, and wearable device data. Update frequencies vary from monthly to a single archival at project completion. Documents contain extensive medical terminology, statistical indicators, and research method descriptions. Fields and units are highly standardized, such as patient ID, diagnostic codes (ICD-10), drug dosage (mg), and observation time (days/months/years), ensuring data traceability and consistency.
Constraints Imposed by These Characteristics on Vector Models and Indexing
The specialized and structured nature of RWE quality documents requires vector models to accurately capture semantic relationships between medical terms and research methods. Inconsistent document update frequencies complicate incremental indexing and periodic full re-indexing strategies, requiring a balance between timeliness and computational resources. Standardized fields and units in documents need special handling during vectorization to avoid generalization or loss of critical information, which would impact recall precision. For example, simple text chunking can fragment crucial dosage-efficacy relationships. Furthermore, the extensive specialized terminology demands pre-trained models with strong performance in the biomedical domain.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk Length | 500–800 characters | RWE documents have a tight logical structure. A moderate chunk length ensures contextual completeness and prevents truncation of critical information. |
Overlap Length | 50–100 characters | Ensures contextual continuity at chunk boundaries, improving recall accuracy for information spanning multiple paragraphs. |
Recall Count | Top 5–8 items | Balances recall precision and processing efficiency, meeting the high information integrity requirements of RWE documents. |
Similarity Threshold | Calibrate empirically | Fine-tune between 0.7–0.85 based on actual recall effectiveness and false positive rates. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | RWE documents may contain numerous charts and complex layouts, requiring sufficient parsing time to avoid timeouts. |
Vector Model | text-embedding-ada-002 or domain-optimized model | Addresses the specialized terminology and complex semantics of RWE documents. Choose a stable general-purpose model or a domain-specific optimized model. |
Three Common Mistakes
- Knowledge base index remains "incomplete" for extended periods, or parts of documents are not indexed: This usually occurs when
PARSE_FILE_TIMEOUT_SECONDSis set too low, causing large or complex documents to time out during parsing and fail to enter the vectorization pipeline. - Key medical terms or data units are ignored in retrieval results: This often happens due to overly coarse chunking strategies that split critical sentences containing specific fields and units, preventing the vector model from accurately capturing their semantics.
- Newly uploaded documents are not retrieved promptly: This might be due to the absence or non-triggering of an incremental indexing mechanism, or a backlog in the
indexing task queue, leading to delays in processing and vectorizing new data.
How to Verify Configuration
- Upload typical RWE document samples and check if the index status for all documents shows "Completed".
- Perform searches using key medical terms and data units from the documents. Verify if the recalled results include relevant context from the original documents and observe the number of recalled items.
- Check the processing speed of the
indexing task queuethrough the FastGPT platform interface to ensure newly uploaded documents are processed promptly. - Define an acceptable recall accuracy and response time. Conduct multiple rounds of testing against this standard, adjusting parameters such as
Similarity Threshold.
Note: The values provided are common starting points. They should be measured against specific samples from the reader's own data.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.