Data Characteristics in Pharmacovigilance
Pharmacovigilance data originates from clinical trial reports, real-world evidence (RWE), spontaneous reporting systems (e.g., China Adverse Drug Reaction Monitoring System, FDA FAERS), and medical literature. Data updates are frequent, typically incremental daily or weekly. Document structures are diverse, including structured case report forms (CRFs), semi-structured free-text descriptions (e.g., adverse event reports), and unstructured medical literature. Core fields include patient demographics, drug information (brand name, generic name, batch number, dosage, usage), adverse event descriptions (terms, onset time, severity, outcome), concomitant medications, and medical history. Units involve dosage (mg, g, IU), frequency (times/day, QID), and time (days, weeks, months), requiring precise parsing.
Constraints Imposed by Data Characteristics on Deployment and Upgrade
High update frequency in pharmacovigilance data requires efficient incremental data ingestion and indexing capabilities to minimize processing latency. Diverse and heterogeneous data structures (structured, semi-structured, unstructured) necessitate flexible data parsers and text processing modules during deployment to accommodate various input formats. Free-text descriptions, in particular, contain numerous medical terms, abbreviations, and colloquialisms, demanding high accuracy from natural language processing models. Data sensitivity (patient privacy) and compliance (GCP, GVP regulations) mandate deployment in private or strictly secure cloud environments, along with fine-grained permission management. Long-term storage and traceability of historical data also impose requirements on storage solutions and data version management, ensuring data integrity during upgrades.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
UPLOAD_FILE_MAX_SIZE | 100 MB | Accommodates documents with extensive medical images or detailed reports, ensuring single-upload completeness. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Handles complex PDFs or OCR-identified reports, preventing parsing failures due to timeouts. |
Chunk size | 800–1200 characters | Balances contextual completeness with model processing efficiency, adapting to adverse event description lengths. |
Recall count | Top 10 entries | Increases recall rate for relevant adverse events or drug information, reducing omission risks. |
Similarity threshold | Calibrate based on actual measurements | Adjusts based on the specificity of medical terminology and the similarity distribution of reports, ensuring accurate matching. |
REINDEX_CRON_EXPRESSION | 0 0 * * * | Triggers full or incremental index updates daily at midnight, adapting to high-frequency data changes. |
Common Pitfalls
- After local deployment, the frontend access address fails to switch from HTTP to HTTPS, leading to browser security warnings or access issues. This occurs due to incorrect SSL certificate configuration on reverse proxy servers like Nginx or Caddy, while internal container services still listen on HTTP ports.
- After private deployment, using the
text2sqltool in workflows to generate SQL fails, reporting syntax errors or database connection exceptions. This happens when thetext2sqlmodel is not optimized for a specific database dialect, or theDATABASE_URLconnection string is configured incorrectly. - Adding a locally deployed language model to FastGPT's model configuration results in an error, stating that the model only supports streaming output. This indicates incorrect configuration of the
MODEL_API_URLorMODEL_API_KEYfor the model service interface, or the model service itself does not support non-streaming request modes.
Verification Steps
- Upload a PDF document containing typical adverse event descriptions. Check if knowledge base segmentation is accurate, if segment length meets expectations, and if no critical information is truncated.
- Search for specific drug names or adverse event terms in the FastGPT knowledge base. Verify that the recall results include all relevant document fragments and check their similarity scores.
- Simulate a data update via the API. Observe system logs for
LOG_LEVELto confirm normal processing, and verify that the updated data is retrievable in the knowledge base.
Note: The values provided are common starting points and should be measured against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.