Data Characteristics
Deviation and Corrective and Preventive Action (CAPA) documents originate primarily from pharmaceutical quality management systems. These include reports on unexpected production events, quality defect analyses, root cause investigations, and CAPA plans. Documents are stored as PDFs, Word files, or scanned images, with PDFs being most common. Document structures typically combine fixed templates and free-form text. Fixed template sections cover structured fields such as deviation numbers, occurrence dates, impact assessments, CAPA action numbers, and completion statuses. Free-form text sections detail deviation phenomena, investigation processes, root causes, and specific implementation steps. Data update frequency correlates with deviation occurrence and CAPA implementation progress, potentially updating multiple times daily. Special fields include batch numbers, product codes, equipment numbers, and various quality parameter units like ppm, mg/L, and ℃.
Constraints Imposed on Knowledge Base Retrieval and Recall
The mixed structure of Deviation and CAPA documents challenges retrieval accuracy. Scanned PDFs require high-quality OCR for effective text extraction and subsequent vectorization. Structured fields like batch numbers and product codes demand precise matching and filtering capabilities from the knowledge base. Free-form text descriptions often contain extensive technical jargon and contextual information, requiring sophisticated segmentation strategies and semantic understanding to prevent loss of critical information or context fragmentation. High update frequency necessitates incremental updates and rapid incorporation of new data into the retrieval scope. The presence of units and specific parameters requires the retrieval system to distinguish between numerical values and units to avoid misjudgments in similarity calculations. For example, a query for "high temperature deviation" should link to text related to the temperature unit "℃".
Configuration Settings
| Configuration Item | Suggested Value | Rationale |
|---|---|---|
Chunk size | 500–800 characters | Balances contextual completeness of free-form text with single-segment processing efficiency, preventing semantic confusion from excessive length. |
Overlap Length | 100–150 characters | Ensures sufficient semantic overlap between paragraphs, reducing information loss, especially when describing complex root causes. |
Similarity threshold | 0.75–0.85 | Deviation and CAPA retrieval demands high accuracy, reducing false positives and preventing irrelevant information from interfering with decision-making. |
Recall count | Top 8–15 entries | Provides enough candidate results for re-ranking models while controlling computational costs. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Handles OCR and text extraction for large scanned PDF documents, preventing processing failures due to timeouts. |
OCR_ENABLED | true | Ensures text content in scanned documents is recognized and included in the knowledge base index. |
Common Pitfalls
- Retrieval results are empty or inaccurate after uploading scanned PDFs. This occurs because
OCR_ENABLEDis not set totrue, preventing text content in scanned documents from being recognized and indexed. - Retrieval of deviation documents for specific batch numbers or equipment numbers yields poor results. This happens when the segmentation strategy is too coarse, leading to critical structured information being truncated or mixed with irrelevant content, impacting precise matching.
- Newly uploaded CAPA documents are not immediately retrievable after a knowledge base update. This typically results from the knowledge base index not being rebuilt promptly or the incremental update mechanism not being configured correctly, preventing new data from being included in the retrieval scope.
Validation Steps
- Upload a deviation report containing scanned images. Use keyword search to verify that OCR-identified text content is complete and accurate.
- Perform precise searches for specific batch numbers and product codes. Check if the returned results include the corresponding documents and validate that structured information within the documents is effectively recalled.
- Upload a new CAPA document. After a reasonable interval, perform a search to confirm the new document is recognized and recalled by the system, verifying the incremental update mechanism functions correctly.
Note: The values provided are common starting points. Measure them against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.