Data Characteristics
Supplier audit documents originate from internal quality management systems, supplier qualification files, on-site audit reports, and historical corrective action records. These data update infrequently, primarily during supplier onboarding, annual evaluations, or significant quality issues. Document structures vary, including PDF contracts, scanned qualification certificates, Word audit report templates, and Excel checklists with defect records. Data fields include supplier name, registered address, production scope, product batch number, audit date, non-conformance descriptions, corrective actions, and completion dates. Some fields may contain specific industry terminology or abbreviations, such as GMP or ISO 13485.
Constraints Imposed by These Characteristics on Deployment and Upgrade
The diverse document formats and relatively static update frequency of supplier audit data create specific requirements for data ingestion and knowledge base construction. During deployment, ensure the file parsing module effectively handles various file types like PDF, Word, and Excel, especially OCR capabilities for scanned PDFs. Infrequent data updates mean the knowledge base's incremental update mechanism needs optimization to avoid reprocessing large amounts of unchanged data. Historical versions must be retained for traceability. Audit reports and qualification files can be large due to numerous images and attachments, requiring the system to handle large files stably. Furthermore, recognizing and indexing industry-specific terminology and standard numbers requires text vectorization models to accurately capture the semantic information of these specialized terms.
Configuration Settings
| Configuration Item | Suggested Value | Rationale |
|---|---|---|
UPLOAD_FILE_MAX_SIZE | 500 MB | Accommodates audit reports and qualification files with many images or attachments |
PARSE_FILE_TIMEOUT_SECONDS | 300 seconds | Ensures large PDF or Word documents have sufficient time to parse |
Chunk size | 800–1000 characters | Balances semantic completeness and recall accuracy, adapting to audit report paragraph structures |
Recall count | Top 8 entries | Covers more potentially relevant information, reducing oversight of critical audit details |
Similarity threshold | 0.75 | Filters for highly relevant audit clauses or non-conformances |
Rerank result count | Top 5 entries | Optimizes final results, focusing on the most important audit document snippets |
Common Pitfalls
- "Offset out of range" errors when uploading large PDF files. This occurs when the file parser encounters memory or index overflow issues with specific large file structures.
- Query results do not include the latest audit corrective actions after a knowledge base update. This happens when the incremental update strategy fails to correctly identify localized modifications in file content.
- After private deployment, the system fails to correctly parse locally stored Excel format checklists. This indicates a missing parsing library or incorrect configuration path.
Verification Steps
- Upload multiple large supplier audit documents in different formats (PDF, Word, Excel). Verify all files are successfully parsed and ingested into the knowledge base.
- Query the ingested audit documents using industry-specific terminology and standard numbers. Check if recall results are accurate and highly relevant.
- Simulate a supplier qualification update scenario by uploading a new version of a document. Confirm the knowledge base correctly identifies and updates relevant information while retaining historical version records.
The values provided are common starting points and should be measured against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.