Data Characteristics
Real-world evidence (RWE) data for clinical trial pre-screening originates from electronic health records (EHRs), insurance claims databases, disease registries, and patient-reported outcomes (PROs). Data update frequencies vary; EHR data may update in real-time, while insurance claims data typically updates in monthly or quarterly batches. Document structures are diverse, including unstructured clinical notes, structured lab results, diagnostic codes (e.g., ICD-10), and medication records. Fields encompass patient demographics, disease diagnoses, treatment plans, adverse events, and follow-up results. Medical terminology, abbreviations, and units of measure (e.g., mg/dL, mmol/L, ℃) are complex and lack standardization.
Constraints on Vector Models and Indexing
The diversity and unstructured nature of RWE data demand high semantic understanding capabilities from vector models. Complex medical terminology and abbreviations require models with strong domain knowledge to accurately capture word meanings. Varying update frequencies across data sources necessitate indexing strategies that efficiently handle incremental updates, avoiding frequent full rebuilds. Diverse document structures, especially the large volume of unstructured text, make traditional keyword-based retrieval inefficient. Vector models are needed to extract deep semantic relationships. Furthermore, non-standardized fields and units increase data preprocessing difficulty, affecting vector representation accuracy and potentially leading to similarity calculation biases.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Embedding Model | text-embedding-v3-large or bge-large-zh-v1.5 | Balances Chinese medical semantic understanding and recall accuracy. |
Chunk Length | 500–800 characters | Balances semantic completeness with vector model input length limits. |
Chunk Overlap Length | 100–150 characters | Ensures contextual continuity and minimizes information loss. |
Recall Count | Top 10–15 | Covers potentially relevant documents, providing sufficient candidates for reranking. |
Similarity Threshold | Calibrate based on measurements | Requires adjustment based on specific datasets and business scenarios to balance recall and precision. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Accommodates parsing large unstructured documents, preventing timeouts. |
Common Pitfalls
- Incorrect model selection leads to semantic understanding deviations. Retrieval results show poor relevance and fail to recognize medical term synonyms or near-synonyms. This occurs because the model lacks medical domain pre-training or has insufficient parameters.
- Unreasonable chunking strategy during knowledge base import. Documents are truncated, and critical information is scattered across different chunks, affecting retrieval completeness. This happens when the length and structural characteristics of RWE documents are not adequately considered.
- OneAPI configuration errors prevent model calls. Logs show connection or authentication failure error codes, such as
401 Unauthorizedor502 Bad Gateway. This indicates incorrect API Key or Endpoint address configuration, or network connectivity issues.
Validation Steps
- Select representative medical query statements and observe whether the recalled results accurately match medical terms and concepts.
- Upload documents containing complex medical reports. Check if chunking fully retains critical diagnostic and treatment information without semantic interruption.
- Use the FastGPT interface to test different Embedding model and chunking parameter combinations. Compare their recall effectiveness in specific clinical scenarios and set qualification thresholds based on business requirements.
- Check OneAPI call status codes in FastGPT logs. Ensure the model interface responds normally without errors like
HTTP 400orHTTP 500.
The values provided are common starting points. Measure performance against your own samples to determine optimal configurations.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.