Data Characteristics
Deviation and Corrective and Preventive Action (CAPA) data originate from reports of abnormal events identified during drug manufacturing, quality control, clinical trials, and post-market surveillance. These event reports are typically structured or semi-structured documents. They include fields such as event description, occurrence date, affected product batches, impact assessment, root cause analysis, corrective actions taken, planned preventive actions, and verification results. Data update frequency aligns with event discovery and processing, potentially daily, weekly, or real-time for urgent events. Document formats vary, including PDF, Word documents, XML, or direct database records. Key fields like "Deviation Number," "CAPA Number," "Occurrence Date," "Closure Date," "Root Cause," and "Impact Level" have clear semantics and data types.
Deployment and Upgrade Constraints from Data Characteristics
The highly structured and semi-structured nature of Deviation and CAPA data requires FastGPT to have flexible document parsing capabilities during deployment to accurately extract key information. Diverse document formats necessitate support for multiple file types for upload and processing. Event-driven update frequencies demand real-time knowledge base updates, requiring consideration of incremental update mechanisms during deployment. Clear field semantics and data types enable precise metadata filtering and retrieval in the knowledge base, reducing irrelevant information. For example, the "Root Cause" field may contain extensive specialized terminology, requiring a high-quality Embedding model for semantic retrieval. Additionally, large volumes of historical data pose challenges for storage and retrieval performance. Data migration and index rebuilding efficiency are critical during upgrades.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
UPLOAD_FILE_MAX_SIZE | 1000 MB | Deviation and CAPA reports may include images or detailed attachments, often resulting in large file sizes. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Parsing large PDF or Word documents can be time-consuming; sufficient timeout is necessary. |
Chunk size (Segment Length) | 800–1200 characters | Ensures each segment contains a complete deviation description, root cause, or action detail, preventing truncation of critical information. |
Rerank result count (Reranked Results Count) | Top 5 entries (Top 5) | Retrieval accuracy is crucial in pharmacovigilance; fine-tuned reranking improves relevance. |
Recall count (Recall Count) | 15–20 entries | Ensures sufficient potentially relevant documents are covered in the initial recall phase, improving recall rate. |
Similarity threshold (Similarity Threshold) | Calibrate based on actual measurements | Test different thresholds to balance recall and precision based on specific datasets and business requirements. |
Common Pitfalls
- File upload fails with "Invalid URL" or "File too large" errors. This may be due to incorrect network configuration of the FastGPT Docker container, preventing access to storage services, or
UPLOAD_FILE_MAX_SIZEbeing set significantly smaller than the actual file size. - Knowledge base Q&A response is slow, especially after reranking. This may be due to a
Chunk size(segment length) that is too small, leading to excessive segments, or selecting a lower-performanceEmbedding model, increasing retrieval and reranking computational burden. - Post-Docker deployment, the PostgreSQL database continuously restarts. This usually indicates improper database configuration parameters, such as insufficient memory allocation, or data volume mounting permission issues, preventing the database from starting and persisting data correctly.
Verification Steps
- Upload multiple Deviation and CAPA reports in different formats (PDF, Word, TXT) and varying sizes (from hundreds of KB to hundreds of MB). Confirm successful file upload and parsing, and that document content displays correctly in the knowledge base.
- Perform knowledge base Q&A for typical deviation queries, such as "root cause for a specific batch product" or "effectiveness of a certain preventive action." Check if the returned results include relevant document snippets and evaluate their accuracy and completeness.
- Check system logs for continuous database restarts, file parsing timeouts, or network connection errors to ensure stable backend service operation.
- Simulate high-concurrency queries to evaluate knowledge base response times, ensuring acceptable performance under actual usage scenarios.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.