Data Characteristics
Quality documents for supplier audits include audit reports, non-conformance lists, Corrective and Preventive Action (CAPA) records, Supplier Quality Agreements (SQAs), and relevant regulatory standards. These documents are typically in PDF, Word, or scanned image formats, with varying degrees of structural organization. Audit reports often have fixed sections like scope, findings, and conclusions, but detailed descriptions can be flexible. CAPA records contain fields such as problem descriptions, root cause analyses, implemented actions, and verification results. Data update frequency is relatively low, usually tied to audit cycles (annual or longer) or when significant non-conformances occur. Documents frequently contain specialized terminology, abbreviations, and unique identifiers like batch numbers and serial numbers from the biomedical field.
Constraints on Knowledge Base Retrieval and Recall
The semi-structured nature of supplier audit documents requires the knowledge base to balance semantic completeness and local detail during chunking. For example, a non-conformance description in an audit report might span multiple paragraphs. Simple sentence-based chunking could lead to information loss. Low update frequency means the knowledge base must process a large volume of historical documents at once, demanding efficient document parsing and embedding. The presence of specialized terminology and abbreviations requires the embedding model to accurately recognize domain-specific concepts, preventing recall failures due to vocabulary mismatches. The specificity of fields and units, such as drug batch numbers, production dates, and specific test value ranges, requires the retrieval system to precisely match or identify approximate values to provide accurate evidence when answering questions like "Did a certain batch of product pass a specific audit?"
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Chunk Length) | 500–800 characters | Balances the completeness of non-conformance descriptions in audit reports, preventing critical information from being cut off. |
Recall count (Recall Count) | 8–12 items | Ensures coverage of multiple potentially relevant audit findings, CAPA records, or SQA clauses, increasing information coverage. |
Similarity threshold (Similarity Threshold) | 0.75–0.85 | Balances recall precision and recall rate, reducing interference from irrelevant documents while not missing highly relevant ones. |
Rerank result count (Reranked Return Count) | 3–5 items | Further filters the initial recall results using a reranking model to select a few high-quality documents that best match the query intent. |
maxContext | 3000 Tokens | Ensures the large language model has sufficient contextual understanding when processing audit reports or CAPA records. |
EMBEDDING_MODEL | text-embedding-ada-002 or domain-specific model | Ensures good understanding and vectorization capabilities for specialized terminology in the biomedical field. |
Common Mistakes
- Knowledge base query results provide generic descriptions instead of specific batch numbers or date information from relevant documents. This happens when chunk granularity is too large or the embedding model's sensitivity to numbers and specific identifiers is insufficient, causing critical factual information to be diluted or ignored during vectorization.
- After uploading a large volume of historical audit documents, knowledge base retrieval response times significantly increase or timeout errors occur. This can be due to
PARSE_FILE_TIMEOUT_SECONDSbeing set too short, preventing the processing of complex or large files, orUPLOAD_FILE_MAX_SIZElimiting file size, leading to some important documents failing to be ingested. - In FastGPT API calls, even with an associated knowledge base, the generated response does not cite document content. This may be because the
kb_idsparameter was not correctly passed during the API call, or theSimilarity threshold(Similarity Threshold) was set too high, leading the model to believe there were no sufficiently relevant documents to cite.
Verification Steps
- For typical audit queries, such as "Query non-conformances for a specific supplier in a particular year," check if the returned citations include corresponding audit reports or CAPA records. Verify the accuracy of key information like batch numbers and dates mentioned within them.
- Upload an audit report containing complex tables and specialized terminology. Observe whether the knowledge base, with the configured
Chunk size(Chunk Length) andRecall count(Recall Count), can correctly chunk and recall paragraphs containing table content or specific terms. - Conduct simulated audit scenario questions, for example, "What are the key clauses of the quality agreement for a specific batch of raw material?" Check if the model's answer accurately cites specific clauses from the Supplier Quality Agreement (SQA) and evaluate if the
Similarity threshold(Similarity Threshold) of the recall results effectively filters relevant content.
The values provided are common starting points and should be measured against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.