Data Characteristics
Real-World Evidence (RWE) data for clinical trial pre-screening originates from Electronic Health Records (EHRs), medical claims databases, registries, wearable devices, and patient-reported outcomes (PROs). This data is highly heterogeneous. It includes structured data like ICD-10 diagnostic codes, lab results, and ATC drug codes. It also includes unstructured data like physician notes and imaging reports. Update frequencies vary; EHRs may update daily, while claims data may update quarterly or annually in batches. Document formats are diverse, including PDF clinical reports, HL7 or FHIR medical records, and CSV or JSON device data. The number of fields is extensive, potentially involving thousands of clinical indicators. Units are complex and inconsistent; for example, blood pressure may be in mmHg, and blood glucose in mg/dL or mmol/L, requiring standardization.
Deployment and Upgrade Constraints
The heterogeneous and large volume of RWE data creates challenges for FastGPT deployment, specifically in data preprocessing and indexing efficiency. Ingesting multi-source, unstructured data requires flexible file parsing and robust data cleaning pipelines. For example, processing large PDF clinical reports requires optimizing the PARSE_FILE_TIMEOUT_SECONDS parameter to prevent parsing timeouts. Frequent data updates, especially from EHRs, demand FastGPT's ability to perform incremental updates and version management to ensure knowledge base timeliness. Inconsistent fields and units necessitate standardization before vectorization. This may involve custom preprocessing scripts or external services, increasing deployment complexity. Additionally, the sensitive nature of medical data requires strict adherence to data security and privacy regulations during deployment, such as setting access controls and data encryption. This may affect model deployment methods and storage locations.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
UPLOAD_FILE_MAX_SIZE | 500 MB | Accommodates large clinical report PDFs and imaging report attachments. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Ensures complete parsing time for complex or scanned PDF files, preventing data loss due to parsing timeouts. |
Chunk size (Segment Length) | 800–1200 characters | Balances semantic integrity of unstructured text with recall efficiency, avoiding excessive truncation of critical clinical information. |
Recall count (Recall Count) | Top 10 | Improves recall rate during pre-screening, covering more potentially relevant clinical features. |
Similarity threshold (Similarity Threshold) | 0.75 | Relaxes the threshold while maintaining relevance to identify more potentially eligible patients from complex medical records. |
Rerank result count (Rerank Return Count) | Top 5 | Optimizes the quality of the final results, prioritizing patient information that best matches clinical trial inclusion/exclusion criteria. |
Common Mistakes
- HTTP 500 errors during document import, with logs showing
Out of memoryorConnection refused. This typically indicates that files are too large or too many concurrent parsing tasks are exhausting server resources or causing parsing service connection timeouts. - Frontend page fails to load, but container logs show normal port listening. This can happen if the
FASTGPT_URLconfiguration does not match the actual access address, or if reverse proxy configuration is incorrect, preventing frontend resources from loading properly. - Query results do not meet expectations, such as missing critical patient information or retrieving irrelevant documents. This may be due to insufficient data preprocessing, such as not standardizing medical terminology, or segmenting strategies that fail to preserve clinical context effectively.
Verification
- Upload RWE documents of various types (PDF, CSV, JSON) and sizes (from a few MB to hundreds of MB). Confirm successful parsing and indexing for all.
- Construct multiple query statements based on typical clinical trial inclusion/exclusion criteria. Verify that FastGPT accurately recalls document segments containing key medical terms, diagnostic codes, and lab indicators. Check the number of recalled items.
- Simulate incremental data updates. Upload new versions of patient data or new medical records. Observe if the new data is retrievable promptly after the knowledge base update and verify its accuracy.
The values provided are common starting points. Measure against specific samples to determine optimal settings.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.