Data Characteristics
Batch record review data primarily originates from paper or electronic batch production records, batch inspection records, and material balance sheets from pharmaceutical manufacturing processes. These records typically exist as scanned PDFs, Word documents, or structured data export files. Data updates align with production batches, usually daily or weekly. Document structures are highly standardized, including fixed headers, tables, signature fields, and date stamps. Examples include production orders, process parameters, equipment usage logs, and environmental monitoring data. Fields include batch number, product name, production date, expiration date, and operator signatures. Units strictly follow pharmacopoeia or industry standards, such as mg, g, L, ℃, and kPa. Table data constitutes a significant portion, recording detailed parameters and deviations for critical production steps.
Constraints on Document Parsing and Chunking
The highly structured nature of batch record review documents, especially the presence of numerous tables, places high demands on document parsing. Standard text chunking methods can truncate table content, losing associations between rows and columns. Handwritten signatures and annotations in scanned PDFs, along with low-resolution scans, increase OCR recognition difficulty, potentially introducing character errors or omissions. Specific terminology and units, such as USP, EP, and GMP, require the parser to accurately identify them to avoid mis-chunking or semantic loss. Batch record continuity and correlation require chunking to preserve the context of adjacent records, allowing subsequent retrieval to obtain complete production process information. A single batch record can contain dozens or even hundreds of pages, posing challenges for single-file processing time and resource consumption.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Chunk Length) | 500–800 characters (characters) | Balances table integrity and contextual relevance, avoiding overly large chunks that dilute information density. |
Chunk overlap (Chunk Overlap) | 50 characters (characters) | Ensures contextual continuity, especially for content spanning pages or tables. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds (seconds) | Addresses the parsing time required for single batch record documents that may contain many pages and complex tables. |
UPLOAD_FILE_MAX_SIZE | 200 MB | Accommodates potentially large scanned files or merged uploads of single batch record files. |
PDF_OCR_ENGINE | PaddleOCR | Optimizes recognition for potential handwritten annotations and complex table structures in batch records. |
Enable Table Recognition | Enabled (On) | Ensures table content is correctly parsed and its structure preserved, preventing data flattening. |
Common Pitfalls
- Some table content in uploaded batch record PDF documents is not chunked, or table rows and columns are disordered. This occurs because the default document parser fails to effectively identify complex table boundaries or tables spanning multiple pages.
- Images (e.g., screenshots of equipment calibration certificates) in imported Word documents do not display in conversations. This happens because image links are not correctly processed or converted to relative paths during document conversion, preventing the frontend from loading them.
- The system frequently reports
VRAM overflowerrors when parsing batch record files. This is due to insufficient GPU memory allocation or improper configuration of the OCR engine when processing high-resolution scans or large batches of files. For example,PDF-markerv0.1may have poor resource management in specific driver environments.
Verification Steps
- Select a batch record PDF document containing complex tables and handwritten signatures. Upload it to the knowledge base. Check if the parsed chunks completely retain the table structure and text information.
- Randomly select 3-5 chunks. Examine their
contentfield to confirm they include key production parameters, batch numbers, and operator information, without obvious truncation or garbling. - Perform retrieval tests using batch record snippets containing specific units (e.g.,
kPa,μg) and terminology. Verify that relevant chunks are accurately recalled and check the completeness of the recalled chunk's context.
The values provided are common starting points. Measure them against your own samples to determine the optimal configuration.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.