Data Characteristics in This Category
Real-world evidence in pharmacovigilance primarily uses data from electronic health records, patient registries, insurance claims databases, and mobile health applications. This data typically combines unstructured text (e.g., clinical notes, adverse event reports) and structured tables (e.g., diagnosis codes ICD-10, medication records ATC). Data update frequencies vary; some high-frequency sources may update daily, while patient follow-up data might update quarterly or annually. Document structures are diverse, including medical reports, laboratory results, and imaging reports. Field names and units can differ across sources; for example, dosage units might be mg or g, and time units might be days or weeks.
Constraints Imposed by These Characteristics on "HTTP Interface and External Systems"
The diversity of real-world evidence data sources requires HTTP interfaces to have flexible data parsing capabilities to accommodate various formats and encodings. Unstructured text content, such as adverse event descriptions, needs advanced text processing and entity recognition to extract key information, increasing data preprocessing complexity. Inconsistent data update frequencies, especially for low-frequency updates, require external systems to configure differentiated pulling strategies to avoid unnecessary resource consumption. Additionally, due to patient privacy concerns, data transmission and storage must comply with regulatory requirements like HIPAA or GDPR, ensuring interface security and compliance. Heterogeneous fields and units mandate standardization and mapping during data ingestion, otherwise affecting the accuracy of subsequent analysis.
Configuration Guidelines
| Configuration Item | Suggested Value | Rationale |
|---|---|---|
HTTP_REQUEST_TIMEOUT | 600 seconds | Real-world data sources are often large, requiring extended transmission and processing times to prevent premature request termination. |
MAX_FILE_SIZE_MB | 500 MB | Supports uploading large structured data files (e.g., CSV, JSONL) and unstructured PDF documents. |
PARSE_PDF_ENABLED | true | Many clinical reports exist in PDF format, requiring PDF content parsing functionality to be enabled. |
CHUNK_SIZE_TOKENS | 800–1200 characters | Balances textual semantic integrity with processing efficiency, accommodating the paragraph length of clinical text. |
RECALL_TOP_K | 5–7 entries | In pharmacovigilance, sufficient relevant context recall is necessary for comprehensive judgment. |
DATA_SOURCE_POLLING_INTERVAL | Calibrate by measurement | Set a reasonable polling interval based on the specific data source's update frequency to avoid frequent requests. |
Three Common Pitfalls
- After uploading files via the HTTP interface, knowledge base content does not increase as expected, or some key information is missing. This often happens because the
file_typeorparse_strategyparameters are not specified correctly, preventing files from being parsed or processed properly. - Frequent
HTTP 504 Gateway Timeouterrors occur when pulling data from external systems. This typically indicates that the data volume is too large, and the processing time for a single request exceeds the default timeout limit of the server or gateway. - Some fields in the data pushed by external systems are empty in FastGPT, or numerical units are inconsistent. This occurs when external system field names do not match the expected field names in FastGPT, or unit conversion and standardization are not performed.
How to Verify Correct Configuration
- Through the FastGPT administration interface, check the file parsing status in the corresponding knowledge base to confirm all uploaded files have been processed successfully.
- Trigger a complete HTTP data pull or push process, then check FastGPT logs for
HTTP 200 OKstatus codes and confirm no timeout errors. - In the FastGPT knowledge base, randomly select several imported real-world evidence data entries. Check that key fields (e.g., adverse event type, drug name, dosage unit) are accurate and consistently formatted.
The values given are common starting points and should be measured against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.