Data Characteristics
Monoclonal antibody (mAb) pharmacovigilance data originates from clinical trial reports, real-world studies, post-market surveillance databases (e.g., FDA Adverse Event Reporting System, FAERS; European Medicines Agency EudraVigilance), and literature. This data exists in both structured (e.g., adverse event report forms in databases) and unstructured formats (e.g., clinician notes, patient interview texts, scientific papers). Structured data includes fields such as patient demographics, drug exposure history, adverse event descriptions (MedDRA coding), severity, and outcome. Unstructured data provides richer narrative details, covering the adverse event process, therapeutic interventions, and outcomes. Data update frequency varies by source; post-market surveillance databases typically update continuously, while clinical trial reports are released at specific times. Documentation includes numerous PDF clinical study reports, package inserts, and review documents, containing detailed descriptions of adverse reactions.
Constraints on Deployment and Upgrade
The diversity and complexity of monoclonal antibody pharmacovigilance data impose specific requirements on FastGPT's deployment and upgrade. Large volumes of unstructured text data (e.g., clinical reports, case narratives) demand efficient text embedding and semantic understanding capabilities. This necessitates attention to the model's long-text processing ability and vector database performance during deployment. Frequent data updates require the knowledge base to support incremental updates, avoiding resource consumption and time delays associated with full rebuilds. The presence of PDF and other document formats challenges the robustness of the file parsing module, requiring accurate text extraction and handling of complex layouts like tables and images. Additionally, adverse event descriptions often contain specialized terminology and abbreviations, requiring high accuracy in tokenization and entity recognition, which impacts knowledge base recall precision and Q&A quality.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
UPLOAD_FILE_MAX_SIZE | 500 MB | Accommodates large PDF documents with charts or scanned content, preventing upload failures. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Monoclonal antibody clinical trial reports are often lengthy, requiring sufficient time for parsing. |
Chunk size | 800–1200 characters | Balances context completeness and embedding model processing efficiency, adapting to the detail level of adverse event descriptions. |
Recall count | Top 8 entries | Ensures coverage of multiple angles and contexts related to adverse reactions, improving recall accuracy. |
Similarity threshold | Calibrated by actual measurement, e.g., 0.75 | Requires testing with actual data and query corpora to balance recall breadth and precision. |
maxContext | 32000 token | Ensures the model maintains context when processing long adverse event narratives and multi-turn queries. |
Common Pitfalls
- After deployment, the knowledge base fails to support continuous follow-up questions, and the model provides irrelevant answers. This usually occurs when
maxContextis set too low, preventing the model from retaining previous conversation history. - Uploading large PDF reports results in timeouts or failures. This may be due to
UPLOAD_FILE_MAX_SIZEorPARSE_FILE_TIMEOUT_SECONDSvalues being insufficient to handle the file size and parsing complexity. - Docker-compose deployment fails to start services, with logs indicating port conflicts or database connection errors. This typically happens when ports are occupied on the host machine or within the Docker network, or database configuration parameters (e.g.,
DB_HOST,DB_PORT) are incorrect.
Verification Steps
- Upload a multi-page PDF clinical trial report for a monoclonal antibody, containing detailed adverse reaction descriptions. Confirm successful file parsing and knowledge base segment generation.
- Ask a complex adverse reaction question related to the uploaded report that requires understanding previous context. Confirm the model can engage in continuous follow-up and provide highly relevant answers.
- Check the FastGPT backend for the number and content of knowledge base segments. Confirm that segment length meets expectations and that key adverse reaction terms and descriptions are accurately extracted.
- Through Docker logs or system monitoring, confirm that all FastGPT container services (e.g.,
fastgpt-api,fastgpt-web,fastgpt-mongo,fastgpt-pg) are running and show no continuous error output.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.