Data Characteristics for This Category
Deviation and Corrective and Preventive Action (CAPA) documents are core records in biopharmaceutical production quality management. Data originates from abnormal event reports on the production floor, quality department investigations and analyses, and subsequent rectification plans and verification reports. These documents typically exist as PDF, Word, or structured XML files. Update frequency is closely tied to production batches and deviation events, with new or updated documents potentially appearing daily. The internal structure of these documents is relatively fixed, including fields such as deviation number, occurrence time, affected product/batch, deviation description, investigation results, root cause analysis, CAPA plan, implementation status, and effectiveness verification. Some fields may contain free-text descriptions, while others are enumerations or dates. Units are typically dates, batch numbers, or specific measurement units (e.g., temperature, pressure).
Constraints Imposed by These Characteristics on "HTTP Interface and External Systems"
The frequent updates and the co-existence of structured and semi-structured data in deviation and CAPA documents impose specific requirements on HTTP interface design and external system integration. First, incremental document uploads and version management must be supported to ensure each update is correctly identified and indexed. Second, given the large amount of free-text descriptions in the documents, the ability to parse text content and extract key information is crucial. The HTTP interface needs to provide file upload capabilities and handle parsing of different file formats. Concurrently, since sensitive quality data may be involved, the interface's authentication and authorization mechanisms must be robust. For specific field formats and units, preprocessing or validation is required during data ingestion to prevent data contamination. Finally, the interface's concurrent processing capability must meet the demands of peak event reporting and document uploads.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
UPLOAD_FILE_MAX_SIZE | 50 MB | Deviation and CAPA documents often include images or scanned copies, potentially leading to large file sizes. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Complex PDF or Word document parsing can be time-consuming; this ensures sufficient time for completion. |
chunk_size | 800–1200 characters | Balances the integrity of long texts with model processing efficiency, reducing semantic loss across chunks. |
overlap_size | 100 characters | Ensures contextual continuity at chunk boundaries, improving recall quality. |
max_connections | 10–20 | Handles concurrent upload requests from external systems; calibrate based on actual measurements. |
api_key_header | X-FastGPT-API-Key | Standard HTTP header naming for easy external system integration and security management. |
Three Common Mistakes
- Error: HTTP interface returns
413 Payload Too Large. Reason: The uploaded deviation document exceeds theUPLOAD_FILE_MAX_SIZElimit. - Error: Key field information is missing or incorrectly identified in the knowledge base. Reason: Document parsing timed out or text extraction rules did not cover specific formats.
- Error: External systems frequently receive
503 Service Unavailablewhen calling the API. Reason: The number of concurrent requests exceeds the processing capacity limit of the FastGPT interface or backend services.
How to Confirm Correct Configuration
- Upload deviation and CAPA documents of different sizes and formats (PDF, Word) to check if they can be successfully uploaded and parsed.
- Randomly select processed documents and query their key fields (e.g., deviation number, CAPA plan) via the FastGPT interface or API to verify information completeness and accuracy.
- Simulate high-concurrency upload scenarios from external systems, observe interface response times and error rates, and ensure service stability under expected load.
- Check system logs for any errors or warning messages related to document parsing or data ingestion.
Note: The values provided are common starting points and should be measured against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.