Data Characteristics for This Category
Deviation and CAPA (Corrective and Preventive Action) registration data primarily originates from internal quality management system documents, production records, laboratory test reports, and related investigation reports. Document update frequency depends on the occurrence of deviation events and the CAPA implementation cycle, typically updated irregularly, but consolidated during annual quality reviews. Document structures are often structured or semi-structured; for example, investigation reports frequently include fixed sections such as event description, root cause analysis, corrective actions, preventive actions, and verification results. Field content involves batch numbers, product names, deviation types, occurrence dates, impact assessments, responsible persons, and completion deadlines. Units include dates, times, quantities, and percentages, often accompanied by extensive free-text descriptions.
Constraints Imposed by These Characteristics on "Document Parsing and Chunking"
The irregular update nature of Deviation and CAPA data requires the document parsing and chunking mechanism to support incremental updates and version management, avoiding reprocessing historical data. Its semi-structured nature and large volume of free text challenge accurate content identification and key information extraction. Traditional parsing methods based on fixed templates may not effectively handle the variability in free-text sections. Entity information within documents, such as batch numbers and product names, requires precise extraction using named entity recognition for subsequent associated queries. Furthermore, the logical relationships between different sections are crucial for understanding the entire deviation event. Therefore, chunking must maintain semantic integrity, preventing the splitting of critical information across different chunks, which could affect subsequent retrieval recall.
Configuration Recommendations
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size | 500–800 Characters | Balances semantic integrity and retrieval accuracy, avoiding overly long or short chunks. |
Overlap Length | 100–150 Characters | Ensures contextual continuity and reduces the risk of information loss. |
Parsing Mode | Smart Parsing | Adapts to semi-structured documents, improving key information recognition rates. |
Image tabletsOCR识别 | Enabled | Ensures key data and descriptions in images can be parsed. |
Scheduled sync interval | weekly | Accommodates irregular document updates, ensuring data timeliness. |
Max File Size | 100 MB | Handles large investigation reports or summary files. |
Common Pitfalls
- Document parsing timeouts, often occurring with large PDF files or scanned documents containing many images, typically due to a low
PARSE_FILE_TIMEOUT_SECONDSconfiguration value. - Question answering results fail to accurately reference image content within documents. This manifests as image links existing in the knowledge base but no response during Q&A, often because the image OCR recognition feature is not enabled or configured.
- When copying parsed text content, a "cannot automatically copy using browser" message appears. This is due to browser security policies restricting automatic copying of cross-origin content.
How to Verify Configuration
- Upload typical Deviation and CAPA documents. Check that the parsed chunks maintain the integrity of key logical units such as event description, root cause, and corrective actions.
- Conduct Q&A tests to verify accurate retrieval and referencing of key field information like batch numbers and product names, as well as the context of free-text descriptions.
- Review parsing logs to ensure no file parsing failures, timeouts, or key information extraction errors occur. Monitor the
OCR Recognition Success Ratemetric. - Simulate adding or updating documents to verify that incremental parsing and knowledge base synchronization function as expected.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.