Data Characteristics for This Category
Quality documents for supplier audits in the biopharmaceutical industry originate from suppliers' quality management system files, production records, inspection reports, change control records, deviation handling reports, and CAPA (Corrective and Preventive Action) documents. These documents are updated according to the supplier's quality management system regulations, such as annual reviews, updates triggered by major changes, or periodic generation based on batch production requirements. Document structures are complex, often in PDF format as scanned images or electronic versions, containing numerous tables, figures, and unstructured text. Fields and units are highly specialized, including batch numbers, expiry dates, production dates, inspection items, test methods, limits, results, and units (e.g., ppm, mg/mL, IU/mg). Cross-references or naming discrepancies may exist across different document types.
Constraints Imposed by These Characteristics on "Document Parsing and Chunking"
The complexity of supplier audit documents demands high accuracy in document parsing. Optical Character Recognition (OCR) accuracy for scanned documents is crucial, especially for those with handwritten annotations or low-quality scans. Correct identification and structured extraction of tables and figures are key, as these often contain core quality data and audit evidence. Accurate recognition of specialized terminology and measurement units directly impacts retrieval precision. The uncertain update frequency requires the system to efficiently handle incremental updates and distinguish between new and old versions. Furthermore, due to cross-references between documents, chunking must preserve contextual relationships to avoid losing critical information through excessive segmentation, which is vital for subsequent audit questioning and traceability.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
UPLOAD_FILE_MAX_SIZE | 500 MB | Accommodates potentially large individual audit reports or system files, preventing upload failures. |
Chunk size (Chunk Length) | 800–1200 characters (characters) | Balances contextual completeness with retrieval efficiency, suitable for long sentences and paragraphs in professional documents. |
Chunk Overlap Length (Chunk Overlap Length) | 150–200 characters (characters) | Ensures continuity of context at chunk boundaries, improving recall of information spanning multiple paragraphs. |
Parsing Strategy | Table Recognition Priority | Table data is a critical source of information in audit documents; prioritizing its structured extraction is essential. |
OCR_ENABLED | True | Addresses the large number of scanned or image-based audit documents, ensuring text content is recognizable. |
MAX_CHUNKS_PER_FILE | Calibrate by actual measurement (Calibrate based on actual measurements) | Limits the number of chunks generated per file, preventing out-of-memory errors or excessively long processing times for very large files. |
Three Common Mistakes
- When uploading Excel files to the knowledge base, the system only recognizes the first two columns, leading to the loss of critical audit data in tables (e.g., batch, result, limit). This occurs because the default parser may not be optimized for multi-column complex tables or
Table Recognition Prioritystrategy is not correctly configured. - Content from uploaded PDF audit reports appears garbled or missing during retrieval, especially in scanned sections. This indicates that
OCR_ENABLEDis not set toTrue, or the OCR engine used performs poorly on this type of document. - After an auditor asks a question, the returned answer lacks critical context, making it impossible to accurately determine the source or completeness of the answer. This may be due to
Chunk size(Chunk Length) being set too short, causing key information to be split across different chunks, or insufficientChunk Overlap Length(Chunk Overlap Length).
How to Verify Configuration
- Upload typical supplier audit documents (including scanned images, complex tables, and multi-page text). Verify that all text content is retrievable after parsing, especially that data in tables is correctly extracted and structured.
- Randomly select specialized terms or key data from documents and perform a knowledge base search. Confirm that the returned chunks contain complete contextual information, without garbled text or omissions.
- Use FastGPT's debugging interface to review parsing logs. Confirm that
OCR_ENABLEDandParsing Strategyconfigurations are effective as expected, and check for any parsing failure or timeout messages.
The values provided are common starting points. They should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.