Data Characteristics
Supplier audit quality documents include audit reports, supplier qualification certificates, Standard Operating Procedure (SOP) documents, inspection records, deviation reports, and change control records. These documents originate from external suppliers. Update frequency depends on the audit cycle and supplier qualification changes, typically annually or biennially. Major changes or issues trigger additional updates. Document structures vary, including standardized templates (e.g., audit reports) and unstructured text (e.g., email correspondence). Fields and units include batch numbers, production dates, expiration dates, test indicators (e.g., percentage content, impurity ppm), equipment models, and serial numbers. Units include milligrams, liters, Celsius, and Pascals.
Constraints on Vector Models and Indexing
Low update frequency for supplier audit documents means less pressure for incremental updates after initial indexing. However, initial indexing requires high throughput. Diverse document structures demand vector models with strong encoding capabilities for various text formats, especially semantic understanding of tables and text within images (after OCR processing). Specific fields and units, such as batch numbers and test indicators, require preserving their semantic relationships during vectorization. This avoids information loss from simple bag-of-words models. Audit documents may also contain numerous regulatory citations and specialized terminology, requiring domain knowledge from the model to ensure accurate and relevant retrieval.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Chunk Length) | 800–1200 characters | Balances contextual completeness for long documents with semantic focus for shorter texts. |
Overlap Length | 100–200 characters | Ensures key information across segments is linked, preventing important content from being truncated. |
Recall count (Recall Count) | Top 10–20 items | Given the strong correlation of audit document content, increasing recall quantity improves relevance coverage. |
Similarity threshold (Similarity Threshold) | Calibrated 0.75–0.85 | Requires testing with specific query needs and document content to ensure high recall. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Accommodates potentially long processing times for large audit reports or OCR of scanned documents. |
UPLOAD_FILE_MAX_SIZE | 500 MB | Addresses the need for a single audit package to contain multiple large files. |
Common Pitfalls
- "Vectorization anomaly" interface error when uploading large knowledge base files, with some chunks failing vectorization. This typically occurs when
PARSE_FILE_TIMEOUT_SECONDSis set too low, causing file parsing or vectorization to time out. - After updating the FastGPT version, existing knowledge bases fail to retrieve effectively, resulting in empty search results. This may be due to changes in the vector model or tokenization strategy in the new version, leading to incompatibility between old indexes and the new model. Re-indexing is required.
- Server crashes after uploading many files, and the knowledge base status remains "Not Ready" after restart. This usually indicates insufficient server resources (e.g., memory or disk space) to handle concurrent indexing tasks, leading to a blocked processing queue.
Validation Steps
- Upload a supplier audit report containing text, tables, and images (processed via OCR). Confirm it is fully parsed and vectorized successfully, without errors.
- Perform multiple queries for key information within the report (e.g., batch numbers, test indicators, regulatory clauses). Verify the accuracy and relevance of the recalled results, and check if the
Recall count(Recall Count) meets expectations. - Simulate uploading multiple large audit documents concurrently. Monitor system resource usage (CPU, memory, disk I/O) to ensure stable operation under high load, without "Not Ready" statuses or timeout errors.
The values provided are common starting points and should be measured against the user's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.