Data Characteristics
Supplier audit regulation data primarily consists of audit reports, Standard Operating Procedures (SOPs), Quality Management System (QMS) documents, regulatory compliance statements, and contract terms. These documents are typically in PDF, DOCX, or scanned image formats. Their structure varies, and they contain extensive specialized terminology, regulation numbers, and technical specifications. The update frequency is relatively stable, usually occurring after regulatory updates, annual audits, or supplier evaluation cycles, rather than high-frequency real-time changes. Documents often include fields such as audit items, defect descriptions, rectification requirements, responsible parties, and completion dates. Date formats are diverse, regulation numbers may combine letters and numbers, and units cover measurements, time, and percentages.
Constraints on Vector Models and Indexing
The heterogeneous nature of supplier audit data requires vector models with strong semantic understanding to accurately capture key information across different document types. For example, the semantic structure of process steps in SOPs differs significantly from defect descriptions in audit reports; a general model might struggle to differentiate them effectively. The need for precise matching of regulation numbers and technical specifications demands high recall accuracy from the index; simple keyword matching is insufficient. The relatively low update frequency means index rebuilding costs are manageable, but each update must ensure the robustness of incremental or full indexing. Furthermore, the quality of Optical Character Recognition (OCR) from scanned documents directly impacts vectorization effectiveness, requiring pre-processing to ensure text content accuracy. Complex tables and nested lists within documents also challenge text splitting strategies, as critical information must not be incorrectly segmented, which would affect vector representation.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Chunk Length) | 500 characters (characters) | Balances contextual completeness with vectorization efficiency, preventing a single vector from carrying too much irrelevant information. |
Chunk overlap (Chunk Overlap) | 100 characters (characters) | Ensures semantic continuity between paragraphs, preventing critical information from being split at paragraph boundaries. |
embedding_model | qwen3-embedding-8b | Optimized for Chinese biomedical terminology, improving vectorization quality for specialized texts. |
maxContext | 32000 token | Ensures the complete context of large audit reports or SOPs can be understood by the model. |
Recall count (Recall Count) | 8 entries (items) | Balances recall breadth with subsequent re-ranking efficiency, covering potentially relevant information. |
Similarity threshold (Similarity Threshold) | Calibrated by actual measurement | Adjusts based on actual query performance and false recall rate, ensuring the quality of recall results. |
Common Pitfalls
- File status remains "Indexing" for an extended period: This usually indicates an incompatible
embedding_modelor insufficient backend resources, preventing the model from loading or running correctly. - Key regulation numbers or technical specifications are missing from query results: This is due to improper text splitting strategies, causing these specific formats to be incorrectly segmented or ignored during vectorization.
- Knowledge base index automatically disappears after some time: This might be related to caching mechanisms or file storage configurations in earlier FastGPT versions, leading to accidental deletion or improper persistence of index files.
Verification Steps
- Upload representative audit reports or SOPs. Check if the file status quickly changes to "Completed" to confirm the
embedding_modelis functioning correctly. - Ask questions specifically targeting regulation numbers or technical specifications within audit reports. Verify if the recall results include this precise information and check the completeness of its context.
- Simulate user queries to evaluate the accuracy and relevance of the Q&A results. Calibrate the
Similarity threshold(Similarity Threshold) based on actual business scenarios. - Regularly check the knowledge base index list in the FastGPT backend to ensure that indexes for uploaded documents persist and are retrievable.
The values provided are common starting points. Measure them against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.