Data Characteristics for This Category
Deviation and Corrective and Preventive Action (CAPA) data primarily originates from internal quality management systems, Manufacturing Execution Systems (MES), and supplier audit reports within pharmaceutical companies. This data typically exists as structured or semi-structured documents, such as deviation investigation reports, CAPA plans, and verification records. Update frequency is high, especially during intensive production batches or early stages of new product launches. Documents often contain extensive technical terminology, flowcharts, standard operating procedure numbers, batch numbers, equipment serial numbers, specific timestamps, and quantitative analysis results (e.g., purity percentages, impurity content in ppm). Key fields include Deviation Number, Occurrence Date, Deviation Level, Root Cause, Corrective Action, Preventive Action, and Verification Status. These documents also include numerous attachments, such as images, charts, and signature pages.
Constraints on Knowledge Base Retrieval and Recall
The specialized and complex nature of deviation and CAPA documents demands high precision in knowledge base retrieval. Highly discriminative fields like timestamps, batch numbers, and equipment serial numbers require accurate identification and indexing to support precise queries. Embedded images and charts often contain critical information; text-only retrieval is insufficient for comprehensive recall. High update frequency necessitates an efficient incremental update mechanism for the knowledge base to avoid data staleness. Furthermore, due to compliance review requirements, traceability of retrieval results is critical. Users must clearly understand which specific document and paragraph sourced an answer. Inter-document relationships (e.g., a CAPA addressing a specific deviation) also require the retrieval system to have some associative recall capability.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk Size | 500-800 characters | Ensures individual knowledge blocks contain sufficient context while avoiding redundancy and reduced vectorization efficiency from excessive length. |
Recall Count | 8-12 items | Balances recall rate with avoiding excessive noise, considering subsequent re-ranking performance. |
Similarity Threshold | 0.75-0.85 | Balances the breadth and precision of recall, reducing the probability of irrelevant documents being retrieved. Requires tuning based on actual data performance. |
Rerank Return Count | 3-5 items | Focuses on core information most relevant to the user, reducing reading burden and providing concise answers. |
UPLOAD_FILE_MAX_SIZE | 500 MB | Accommodates large files that may contain numerous embedded images and attachments in deviation reports, ensuring successful uploads. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Provides sufficient time to process complex PDFs or Word documents with many charts, preventing parsing timeouts. |
Common Pitfalls
- Query results lack image or chart information, requiring users to locate original documents separately. This occurs when file parsers fail to extract non-text content from original documents or when the retrieval system is not designed to index image content.
- When users query deviations related to specific batches or equipment, the system fails to recall relevant documents and returns generic answers. This happens when the knowledge base indexing does not adequately identify and utilize structured metadata such as
Batch NumberorEquipment Serial Number. - During knowledge base Q&A, the system cannot pinpoint the specific source document and paragraph for an answer, leading users to doubt the answer's reliability. This occurs when the retrieval system does not implement or enable document traceability, failing to link answers to original knowledge blocks.
How to Verify Configuration
- Select a deviation report containing critical charts and complex technical terms. Conduct multiple rounds of questioning to check if answers accurately cite key information from the report and provide corresponding document source links.
- Use queries that include specific batch numbers, equipment IDs, and other metadata. Verify if recall results precisely match relevant reports. Observe if
Recall CountandSimilarity Thresholdmeet expectations. - Simulate uploading a new CAPA document. Check if the knowledge base completes indexing within a reasonable timeframe and if simple queries can verify that new content is retrievable. Monitor file processing status codes for normalcy.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.