Data Characteristics
Deviation and Corrective Action and Preventive Action (CAPA) documents originate from quality management systems in biopharmaceutical manufacturing. Data sources include batch records, quality control reports, audit reports, and equipment calibration records. These documents are typically PDF scans, Word documents, or structured database exports. Update frequency is high, with new deviation events and CAPA measures generated continuously. Document structures usually include fixed fields such as event description, root cause analysis, corrective actions, preventive actions, responsible person, completion date, and status. However, specific formats vary by enterprise and system. Field content is often unstructured text descriptions. Units include time (e.g., hours, days), quantity (e.g., batches, units), and temperature (e.g., Celsius).
Constraints on Deployment and Upgrade
Frequent updates and diverse formats of Deviation and CAPA documents require a deployment solution with high availability and flexibility. The documents contain extensive specialized terminology and abbreviations, challenging model comprehension. This necessitates configuring dedicated glossaries or domain-specific knowledge bases. The sensitive nature of document content dictates the importance of data isolation and access control, requiring data security and compliance during deployment. The prevalence of unstructured text means document parsing can be time-consuming, demanding significant computing resources and higher timeout settings. Furthermore, integrating with numerous document generation systems requires considering various data interfaces and conversion mechanisms. Upgrades must ensure compatibility and smooth data migration.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
UPLOAD_FILE_MAX_SIZE | 500 MB | CAPA documents may contain numerous images or attachments, ensuring large file uploads are not restricted. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | OCR and structural analysis of complex PDF scans can be time-consuming, preventing parsing timeouts. |
maxContext | 8000 tokens | Deviation descriptions and root cause analysis often have large text volumes, ensuring full context understanding. |
Chunk size | 500 characters | Ensures each text segment contains sufficient semantic information while avoiding excessive length that impacts recall. |
Recall count | Top 10 entries | CAPA documents have strong interconnections; increasing recall items improves relevant information coverage. |
Similarity threshold | 0.78 | Precisely matches specialized terms and key information, reducing irrelevant results. |
Common Pitfalls
- Files remain in a "processing" state for an extended period after upload, eventually failing with a "parsing timeout" error. This usually occurs when the
PARSE_FILE_TIMEOUT_SECONDSparameter is set too low, failing to accommodate the parsing time of complex documents. - After a system upgrade, the FastGPT container fails to start, and logs show "MongoDB connection refused." This might stem from a mismatch in the MongoDB service's network configuration within the
docker-compose.ymlfile during the upgrade, leading to inter-service communication failure. - After deployment, when the number of concurrent users exceeds a certain threshold, some requests remain in a waiting state for a long time. This typically indicates insufficient server resources (CPU, memory) to support the anticipated concurrent load.
Verification Steps
- Upload a CAPA PDF scan containing images and complex tables. Confirm successful parsing and reasonable content segmentation.
- After starting the FastGPT service, check the logs of all dependent services (e.g., PostgreSQL, MongoDB, Ollama). Confirm no abnormal startup errors.
- Simulate multiple users simultaneously querying uploaded Deviation and CAPA documents. Observe response times and compare them against defined performance metrics.
- Review the parsed document content. Confirm the accuracy of key field extraction (e.g., "corrective actions," "root cause"). Compare with manual review results to calibrate recall and similarity thresholds.
The values provided are common starting points and should be measured against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.