Data Characteristics
GMP compliance data originates from regulatory documents published by national and international pharmaceutical regulatory agencies, guidelines from industry associations, and internal quality management system documents (e.g., SOPs, batch production records, validation reports). These documents exist as PDFs, Word files, or internal knowledge management system pages. Regulatory documents update infrequently, typically every few years. However, revised regulations have a broad impact. Internal SOPs and guidelines update based on changes in production processes, equipment introductions, or regulatory adjustments, potentially monthly or quarterly. Document structures typically include strict chapter numbering, clause content, revision history, and attachments. Fields often involve batch numbers, effective dates, revision numbers, and document numbers. Some documents also contain flowcharts and tabular data.
Constraints from "Citing Sources and Traceability"
The authoritative and rigorous nature of GMP compliance documents requires the AI Agent to provide precise citations for compliance judgments. The revision history of regulations and SOPs makes version traceability critical. This ensures that the current, effective version is cited, or a specific historical version is cited based on the query timestamp. The extensive use of specialized terms and acronyms in documents requires the knowledge base to accurately identify and associate context. Furthermore, metadata such as document numbers and effective dates are crucial for determining document applicability and require focused attention during retrieval and citation. When knowledge base content has low relevance to a user's query, irrelevant content should not be returned. This requires retrieval strategies with high precision and low retrieval redundancy.
Configuration Settings
| Configuration Item | Suggested Value | Rationale |
|---|---|---|
Chunk size (Chunk Size) | 500-800 characters | GMP document clauses are often long. This avoids excessive splitting that leads to context loss while maintaining vector retrieval efficiency. |
Chunk Overlap Length (Chunk Overlap) | 50-100 characters | Ensures information at chunk boundaries is not lost, improving retrieval recall. |
Recall count (Retrieval Count) | Top 5-8 items | Given the rigor of the documents, retrieving more relevant clauses helps the model make comprehensive judgments. |
Similarity threshold (Similarity Threshold) | 0.75-0.85 | Ensures the precision of retrieved content, preventing irrelevant regulatory clauses from being cited. |
Rerank result count (Reranked Return Count) | Top 3 items | After reranking, focusing on the few most relevant items reduces the model's processing burden and improves answer accuracy. |
Metadata Filtering | Effective Date >= Current Date (Effective Date >= Current Date) | Ensures only the latest effective versions of regulations and SOPs are cited. |
Common Pitfalls
- The model returns citation links pointing to internal Notion links that users cannot directly access. This occurs because the knowledge base did not convert Notion links to publicly accessible URLs during import.
- When a user asks about compliance issues for a specific batch number, the model fails to cite relevant batch production records. This happens because the knowledge base did not effectively parse and index tabular data within the documents.
- When a user queries in Chinese, the model cannot cite English GMP documents from the knowledge base. This is due to the tokenizer or vector model not being optimized for mixed-language content, leading to semantic mismatch.
Verification
- Select several typical compliance questions. Test whether the model's answers provide citations precise down to the clause level and verify that citation links are directly accessible.
- Randomly select expired old versions of SOPs or regulations. Query the model about related content to verify that the model avoids citing these outdated documents through the
Metadata Filteringmechanism. - Submit queries containing specific document numbers, batch numbers, or effective dates. Cross-reference whether the model's returned citations accurately match this metadata information.
- Compare the model's citation performance for mixed Chinese and English queries. Ensure the knowledge base can effectively retrieve and cite relevant documents in a multilingual environment.
The values provided are common starting points and should be measured against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.