Document Parsing and Chunking for Deviation and CAPA Products

Deviation and CAPA (Corrective Action and Preventive Action) product data originates primarily from internal quality management systems, Manufacturing

Data Characteristics

Deviation and CAPA (Corrective Action and Preventive Action) product data originates primarily from internal quality management systems, Manufacturing Execution Systems (MES), and Laboratory Information Management Systems (LIMS). These systems generate numerous reports, records, and process documents. Document structures are typically highly standardized, including deviation reports, investigation reports, Root Cause Analysis (RCA), and CAPA plans and execution records. Field content covers event descriptions, occurrence times, responsible parties, impact assessments, corrective actions, preventive actions, and verification results. Documents often include specific identifiers such as product batch numbers, equipment IDs, and SOP (Standard Operating Procedure) version numbers. Document update frequency is relatively low; archiving usually occurs after a deviation, investigation completion, or CAPA action execution and verification.

Constraints on Document Parsing and Chunking

The standardized structure of Deviation and CAPA documents requires accurate identification and extraction of key information during parsing, such as event type, root cause, and CAPA status. Batch numbers and SOP version numbers are critical for precise matching during retrieval. Chunking must specifically address these identifiers to prevent truncation or confusion with other information. Due to the low update frequency of these documents, real-time requirements are not stringent. However, historical data completeness and accuracy are paramount, necessitating proper handling of all historical document versions during parsing. Documents often contain extensive tabular data, such as CAPA action lists and verification results. Parsers must effectively identify and process table structures to prevent data loss or misalignment.

Configuration Guidelines

Configuration ItemRecommended ValueRationale
Chunk size (Chunk Size)800–1200 charactersBalances contextual completeness with retrieval efficiency, suitable for report-style documents.
Chunk Overlap100–200 charactersEnsures critical information overlaps between adjacent chunks to prevent semantic fragmentation.
UPLOAD_FILE_MAX_SIZE500 MBAccommodates large deviation investigation reports and attached files.
Parsing StrategyBy TitleDeviation and CAPA documents typically have clear chapter title structures.
Table HandlingEmbed as TextConverts table content into continuous text to maintain data integrity.
PARSE_FILE_TIMEOUT_SECONDS600 secondsProvides sufficient time to process complex PDFs or documents containing numerous images.

Common Pitfalls

  • Uploading large PDF documents results in a 413 Request Entity Too Large error. This occurs when Nginx or the web server's request body size limit is lower than UPLOAD_FILE_MAX_SIZE.
  • Key batch numbers or equipment IDs are incomplete or missing in retrieved results after parsing. This happens when these identifiers are truncated during chunking, leading to semantic loss.
  • Multi-level directory content from Feishu or online document platforms is not fully parsed, making some information unretrievable. This is due to the connector failing to recursively traverse all subdirectories or incorrectly handling the specific format of online documents.

Verification Steps

  • Upload a deviation report PDF containing tables and multiple chapters. Verify that the parsed chunks completely retain table content and chapter semantics.
  • Upload a CAPA record with key identifiers such as batch numbers and SOP version numbers. Use keyword search to confirm precise retrieval of these identifiers.
  • Select an online document (e.g., a Feishu document) with a multi-level directory structure. Verify that all levels of its content are parsed and retrievable.

The values provided are common starting points and should be measured against the reader's own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.