Data Characteristics of This Category
Supplier audit quality documents include audit reports, non-conformance lists, Corrective and Preventive Action (CAPA) records, supplier qualification certificates, on-site photos, and related communication records. These documents originate from supplier submissions, audit team on-site records, and internal quality management systems. Document update frequencies vary. Audit reports are typically periodic (e.g., annual or biennial), while CAPA records update in real-time based on non-conformance resolution progress.
Audit reports are often structured or semi-structured, containing clear section titles and data tables. Non-conformance lists are presented in a list format, with fields such as non-conformance description, responsible party, planned completion date, and actual completion date. Field units typically involve dates, percentages (e.g., defect rates), and quantities. Text descriptions may contain extensive specialized terminology and abbreviations.
Constraints Imposed by These Characteristics on Model Integration and Configuration
The semi-structured nature of supplier audit documents requires models to handle both structured information extraction and unstructured text understanding. Extracting table data from audit reports is crucial, demanding models with table parsing capabilities to ensure correct field-value associations.
CAPA records have a high update frequency. Knowledge base synchronization mechanisms require flexible configuration, favoring incremental updates over full rebuilds. Specialized terminology and abbreviations in documents necessitate strong domain adaptability during embedding and retrieval. This may require fine-tuning or enhancement with domain-specific data. The presence of numerous dates and percentages highlights the need for data format validation after information extraction to prevent downstream application parsing failures due to incorrect model output formats.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Segment Length) | 800–1200 characters | Balances semantic completeness of long texts with retrieval efficiency of short texts, preventing truncation of key information in audit reports. |
Overlap Length | 100 characters | Ensures contextual continuity and handles logical connections across paragraphs in audit reports. |
Similarity threshold (Similarity Threshold) | 0.75 | Improves retrieval accuracy, reduces interference from irrelevant documents, and applies to regulatory compliance-related queries. |
Recall count (Number of Retrieved Items) | Top 5 | Balances model processing capability with information completeness, ensuring the model obtains sufficient context for decision-making. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Handles large audit reports or documents containing many images/tables, preventing parsing timeouts. |
maxContext | 8192 | Accommodates long text input models, providing richer contextual information when addressing complex audit issues. |
Common Mistakes
- Model returns an empty thought process or chaotic output logic. This occurs when the
temperatureparameter is too high ortop_pis improperly configured, leading to overly divergent answer generation that fails to meet the strict requirements of audit documentation. - File parsing fails or stalls after uploading large supplier audit reports. This happens when
UPLOAD_FILE_MAX_SIZEorPARSE_FILE_TIMEOUT_SECONDSparameters are set too low, meaning the file exceeds system processing capacity or parsing time limits. - Model testing succeeds on the model provider's page but returns a
404error after adding the model in FastGPT. This is due to an incorrectmodelIdor a mismatch between the model routing configuration in OneAPI and the name expected by FastGPT, preventing requests from being correctly forwarded.
How to Verify Configuration
- Upload a representative supplier audit report. Check if the knowledge base segment preview is complete and free of obvious semantic truncation.
- Ask the model questions about specific non-conformances or CAPA records in the audit report. Verify if the model's answer accurately quotes the original text and if the cited document snippets are correct.
- Simulate complex audit query scenarios, such as cross-document information integration. Observe if the model can effectively retrieve and synthesize multiple document snippets to provide coherent and logically correct answers. Set reasonable thresholds based on expert feedback.
Note: The values provided are common starting points. Measure them against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.