Data Characteristics
Deviation and Corrective and Preventive Action (CAPA) data originates from internal record systems within drug manufacturing, quality control, and post-market adverse event reporting processes. This data typically exists in structured or semi-structured documents, such as PDFs, Word documents, XML, or JSON files. Update frequency depends on the real-time requirements for event occurrence and processing. High-risk deviations may require immediate recording and handling, while low-risk deviations or CAPA updates have defined cycles. Document structures commonly include fields like event description, occurrence time, affected product batch, scope of impact, investigation results, root cause analysis, corrective actions, preventive actions, responsible parties, and completion deadlines. Fields may involve specific coding systems, such as drug batch numbers, product codes, ICD-10 codes (for adverse event classification), timestamps (RFC 3339 format), and numerical risk scores.
Constraints Imposed by "HTTP Interface and External Systems"
The highly structured and real-time nature of Deviation and CAPA data requires careful attention to accurate data model mapping and transmission efficiency during HTTP interface integration. Diverse document formats necessitate robust file parsing capabilities to extract key information. For example, text fields containing ICD-10 codes require correct identification and parsing to prevent data loss or misinterpretation. Timestamp fields, such as event occurrence time, must strictly adhere to a unified time format standard, such as ISO 8601 or RFC 3339, to ensure consistent time synchronization across systems. Additionally, given the sensitive quality and safety information involved, interfaces need to support high concurrent access and data integrity validation mechanisms. Precise error response and traceability are crucial during inter-system interactions. For instance, when an external system returns a non-200 status code, clear error messages and retry strategies are necessary.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
maxContext | 3000 tokens | Balances complex deviation descriptions with performance, preventing single request overload. |
Chunk size (Segment Length) | 500 characters (characters) | Ensures semantic integrity after splitting long texts, facilitating indexing and recall. |
PARSE_FILE_TIMEOUT_SECONDS | 120 seconds (seconds) | Most deviation and CAPA document parsing completes within this timeframe, preventing prolonged blocking. |
Similarity threshold (Similarity Threshold) | 0.75 | Higher than typical thresholds, ensuring recalled CAPA suggestions are highly relevant to the current deviation. |
HTTP_REQUEST_TIMEOUT_MS | 30000 ms (milliseconds) | Accommodates potential response delays from external systems, such as large database queries or complex logic processing. |
Retry Interval | 5 seconds (seconds) | Provides a reasonable retry interval for temporary external system errors, such as network fluctuations or service restarts. |
Common Pitfalls
- HTTP requests return 400 or 500 status codes, but the response body lacks specific error information. This makes it difficult to determine if the issue is an incorrect request parameter format or an internal error in the external service.
- Imported documents remain in an "indexing" state for extended periods. This occurs when documents contain a large amount of non-text content or complex table structures, exceeding the default parser's processing capabilities.
- Data fields returned by the external system are empty, such as
CAPA_IDorROOT_CAUSE. This typically indicates inaccurate field mapping in the interface configuration or missing data in the external system itself.
Validation Steps
- Perform import operations on at least 5 typical Deviation and CAPA documents. Verify their indexing status is "completed."
- Select representative deviation data. Query it via the HTTP interface and cross-reference key fields in the returned results, such as
CAPA_ID,DEVIATION_TYPE, andEFFECTIVE_DATE, to ensure consistency with the original data. - Simulate scenarios where the external system returns non-200 status codes and specific error messages. Verify that FastGPT correctly captures and displays errors, and that the predefined retry logic is triggered.
- Use FastGPT's retrieval function to query CAPA records related to imported deviations. Check if the
similaritymetric of the recalled results meets business requirements to validate theSimilarity threshold(similarity threshold).
Values provided are common starting points and should be measured against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.