Data Characteristics
Deviation and Corrective and Preventive Action (CAPA) documents are central to quality management systems in the biopharmaceutical industry. These documents typically exist as PDFs, Word files, or structured database records. Data sources include anomaly records from production processes, laboratory test result deviations, audit findings, customer complaints, and supplier quality issue reports.
Deviation records are often real-time or near real-time, generated immediately upon an anomaly. CAPA document creation, execution, and verification cycles can range from weeks to months.
Document structures commonly include fields such as Deviation Number, Date of Occurrence, Description, Root Cause Analysis, Corrective Actions, Preventive Actions, Responsible Person, Completion Date, and Verification Results. Units involved include batch numbers, instrument serial numbers, measurement values and their units (e.g., mg/mL, pH, °C), and time units (hours, days).
Constraints on Knowledge Base Retrieval
The real-time nature of deviation and CAPA documents requires the knowledge base to have low-latency data synchronization. This ensures retrieval of the latest status of actions.
These documents contain extensive specialized terminology, abbreviations, and specific identifiers (e.g., GMP-001, SOP-QC-005). This challenges the domain adaptability of embedding models, requiring accurate contextual understanding.
The multi-field structure means queries may involve combined retrieval across multiple dimensions. For example, finding "CAPA actions for a specific batch under a particular temperature deviation."
Documents often include tables and images. Pure text knowledge base chunking may lose critical information, necessitating consideration of multimodal or enhanced text extraction.
CAPA action verification results are important for decision-making. Retrieval must identify and prioritize verified or in-progress actions, avoiding recommendations for outdated or invalid solutions.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk Length | 500–800 characters | Ensures each knowledge chunk contains a complete deviation description or CAPA action detail. Avoids excessive length that leads to information redundancy and reduced embedding efficiency. |
Recall Count | 8–12 items | Deviation and CAPA issues may involve multiple related historical cases or actions. Increasing recall count improves coverage. |
Similarity Threshold | 0.75–0.85 | The domain has many specialized terms. A threshold that is too low may recall irrelevant content, while one that is too high may miss semantically similar but differently worded documents. |
Rerank Return Count | 5 items | Reranks recall results to prioritize CAPA actions that best match the query intent or are most recently updated, improving accuracy. |
Embedding Model | text-embedding-ada-002 or domain-fine-tuned model | text-embedding-ada-002 offers good baseline performance. A model fine-tuned on biopharmaceutical data can improve understanding of specialized terminology if available. |
API Request Timeout | 60 seconds | Allows for processing large amounts of text or complex queries. Extends timeout to prevent request failures due to network or processing delays. |
Common Mistakes
- Symptom: Knowledge base retrieval results do not include deviation or CAPA documents clearly related to the query. Reason: The
Similarity Thresholdis set too high, leading to overly strict recall and excluding potentially relevant documents. - Symptom: When calling the knowledge base API, the response does not cite relevant documents from the knowledge base, or the cited content has weak relevance to the question. Reason:
maxContextor thepromptinsufficiently emphasizes cited knowledge, or theRecall Countis too low, preventing the model from obtaining enough context for effective citation. - Symptom: The system prompts "No available Embedding model." Reason: The specified embedding model, such as
text-embedding-ada-002, is not correctly configured or enabled in the configuration file or OneAPI interface.
Verification
- Execute a series of queries containing biopharmaceutical specialized terminology and deviation/CAPA-related content. Examine the
Similarityscore distribution of the recalled results and manually evaluate the relevance of the top few results. - For typical deviation scenarios, query FastGPT and verify whether the knowledge points cited in the answer accurately correspond to the expected deviation or CAPA document content.
- Conduct batch query tests via the API interface. Monitor
API Request Timeoutto ensure no significant request failures occur under expected load.
The values provided are common starting points and should be measured against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.