Data Characteristics in this Category
Pharmacovigilance registration dossiers primarily use adverse event reports, safety update reports, risk management plans from marketing authorization holders, and global medical literature database search results. Data updates frequently. After drug launch, adverse event reports continuously flow in, and safety update reports are typically submitted quarterly or annually. Documents vary in format, including structured database records (e.g., ICH E2B), semi-structured PDF reports (e.g., CIOMS I forms), and unstructured full-text medical literature. Fields cover patient demographics, drug information, adverse event descriptions (including MedDRA coding), causality assessments, and follow-up actions. Units include dosage (mg, g, IU), frequency (times/day, week), and time (days, months, years), often with medical terminology and abbreviations.
Constraints on "Deployment and Upgrades"
The high update frequency of pharmacovigilance data requires FastGPT's knowledge base to support efficient incremental updates and version management, ensuring the timeliness of submission documents. Diverse document formats necessitate robust document parsing tools that can handle mixed structured and unstructured data inputs and accurately extract key information. The complex semantics and specialized terminology in medical literature demand high model comprehension. Given sensitive patient information, data security and privacy are critical deployment considerations, requiring strict access controls and data anonymization mechanisms. Standardizing fields and units is fundamental for accurate RAG recall and precise question answering; any parsing or indexing deviation can lead to errors in submission documents.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
UPLOAD_FILE_MAX_SIZE | 500 MB | Pharmacovigilance reports and literature often contain numerous images and charts, leading to large file sizes. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Processing complex PDF reports and lengthy medical literature requires extended parsing times. |
maxContext | 3000 Tokens | Ensures capture of complete context and detailed descriptions within adverse event reports. |
Chunk size (Segment Length) | 800–1200 characters | Balances context completeness and recall efficiency, preventing loss of information in long paragraphs. |
Similarity threshold (Similarity Threshold) | 0.75 | Ensures high relevance of recall results to query intent, reducing interference from inaccurate information. |
Rerank result count (Reranked Return Count) | Top 10 | Improves the hit rate for key information, covering more relevant adverse event reports. |
Three Common Mistakes
- A provided link fails to parse document content because the document parsing service is not running or the configured proxy cannot access external resources.
- The online version of the chatbot continuously searches the knowledge base without responding. This usually indicates an overloaded knowledge base indexing service or abnormal database connection.
- Parsing accuracy decreases after a version upgrade. This happens when the new model or parser is incompatible with existing data formats, requiring adjustments to parsing rules or re-indexing.
How to Verify Correct Configuration
- Upload a typical adverse event report PDF and verify that it parses correctly and extracts key fields such as patient age, drug name, and adverse event MedDRA codes.
- Perform an incremental knowledge base synchronization for a recently updated safety report. Confirm that new data is successfully indexed and included in recall.
- Use queries containing medical terminology to test FastGPT's question-answering accuracy for pharmacovigilance-related questions. Cross-reference the recalled source documents with the answers for relevance.
- Monitor system logs to confirm no abnormal errors or prolonged timeouts occur during file upload, parsing, vectorization, and knowledge base querying.
Note: The values provided are common starting points. Measure against your own samples for optimal performance.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.