Data Characteristics for this Category
Deviation and Corrective and Preventive Action (CAPA) documents are central to biopharmaceutical production quality management. Data originates primarily from Quality Management Systems (QMS), Manufacturing Execution Systems (MES), and Laboratory Information Management Systems (LIMS). Documents are typically structured or semi-structured reports. They contain fields such as event descriptions, root cause analyses, impact assessments, corrective actions, preventive actions, verification results, and closure statuses. Update frequency varies based on deviation severity and CAPA execution cycles. High-risk deviation CAPAs may require completion within days, while complex CAPA verification and closure can extend over several months. Documents often include text descriptions, images (e.g., equipment photos, batch record screenshots), and associated identifiers like batch numbers, product codes, equipment IDs, and SOP version numbers.
Constraints Imposed by These Characteristics on "Referencing and Tracing"
The structured and semi-structured nature of Deviation and CAPA documents requires the knowledge base to effectively identify and extract key fields during indexing, such as root causes, action contents, and verification results, to support precise retrieval. Their update frequency and lifecycle management necessitate incremental update and version management capabilities within the knowledge base, ensuring the real-time accuracy of referenced content. The inclusion of images and associated identifiers in documents demands multimodal processing capabilities and entity recognition from the RAG (Retrieval Augmented Generation) system. This ensures accurate referencing of image descriptions or linking to external data when generating responses. Furthermore, the stringent requirements for referencing in registration and declaration documents mean that the tracing mechanism must precisely point to specific paragraphs or even specific figures within original documents to meet compliance review needs.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size | 500–800 characters | Balances semantic completeness and retrieval efficiency, preventing overly long segments from diluting key information. |
Recall count | 8–12 entries | Ensures coverage of multiple potentially relevant deviation or CAPA cases, increasing information richness. |
Similarity threshold | 0.75–0.85 | Strictly filters highly relevant document fragments, reducing interference from irrelevant information. |
Rerank result count | 4–6 entries | Focuses on the most relevant pieces of information, reducing model processing load and improving response quality. |
maxContext | 3000–4000 token | Accommodates sufficient retrieved content while reserving space for model response generation. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Addresses the parsing needs of large deviation reports or CAPA documents containing multiple images. |
Three Common Pitfalls
- Knowledge base retrieval results are empty, but keyword searches locate relevant documents. This may occur if text segmentation granularity is too coarse or if indexing strategies do not adequately consider key entities within documents.
- AI responses cite CAPA numbers or SOP version numbers that do not match original documents. This typically happens when the knowledge base fails to correctly extract or recognize structured fields within documents.
- During response generation, the citation source only displays the document title, without pinpointing specific paragraphs. This occurs when document parsing does not retain sufficient metadata or lacks a paragraph-level tracing mechanism.
How to Verify Correct Configuration
- For typical deviation events, ask for suggestions to generate relevant CAPAs. Check if the CAPA numbers and action contents cited in the response are accurate and traceable to specific locations in the original document.
- Upload CAPA reports containing images and tables. Ask specific questions about image content or table data. Verify if the AI response accurately describes and references the corresponding images or tables.
- Simulate historical deviation cases and ask for their root causes. Check if the AI response accurately references the root cause analysis section in historical documents and provides precise citation links.
The values given are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.