Data Characteristics for this Category
Deviation and CAPA (Corrective and Preventive Action) system data originates primarily from internal quality management systems, Manufacturing Execution Systems (MES), and document management systems. Update frequency often correlates with production batches and quality event processing, showing irregular but periodic patterns (e.g., weekly, monthly, or event-triggered). Document structures typically combine structured and unstructured data, such as deviation reports, root cause analysis reports, CAPA plans and execution records, and verification reports. These documents often include specific fields like deviation number, occurrence time, affected product batch number, deviation description, impact assessment, root cause, CAPA measures, responsible person, completion deadline, and verification results. Units commonly involve time (hours, days), quantity (batches, pieces), and temperature (°C), with high demands for data precision.
Constraints Imposed by These Characteristics on "Vector Model and Indexing"
The mixed structure of deviation and CAPA documents requires vector models to effectively integrate long text descriptions with key structured fields. Irregular update frequencies mean the index must support incremental updates and version management to ensure the knowledge base's timeliness. Field specificity, such as deviation numbers and batch numbers, demands that the indexing process identifies and assigns higher weight to these critical entities to improve recall precision. Furthermore, a strong correlation exists between CAPA measures and verification results, requiring the vector model to capture this causal or dependency relationship during semantic matching. The demand for data precision necessitates maintaining contextual completeness during chunking to avoid losing critical information or introducing semantic deviations due to misinterpretation.
Configuration Settings
| Configuration Item | Recommended Value | Rationale for this Value |
|---|---|---|
Chunk Size | 800–1200 characters | Deviation and CAPA reports often contain lengthy descriptive text and analysis processes. Longer chunks help capture complete semantics and prevent critical information from being truncated. |
Overlap Size | 100–200 characters | Ensures sufficient contextual connection between adjacent chunks, improving the coherence of cross-chunk semantic recall. |
Similarity Threshold | 0.75–0.85 | Deviation and CAPA Q&A demands high precision, requiring a higher similarity to reduce the recall of irrelevant content. |
Recall Count | Top 5 | Given the specialized and targeted nature of Q&A, recalling a small number of highly relevant document snippets effectively controls model input length and improves response efficiency. |
Embedding Model | text-embedding-ada-002 or Doubao-embedding | Selects mainstream embedding models with strong semantic understanding and domain adaptability to effectively process specialized terminology and complex logic. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Deviation and CAPA reports may contain numerous attachments or complex formats. Extending the file parsing timeout ensures complete file processing. |
Three Common Pitfalls
- The knowledge base returns irrelevant results, indicated by recalled document snippets semantically deviating from the query intent. This occurs when the
Similarity Thresholdis set too low, leading to the recall of overly generalized content. - The model cannot accurately answer questions involving specific deviation numbers or batch numbers, indicated by the absence of critical entity information in the answer. This happens when the document chunking strategy fails to effectively preserve or emphasize the context of these structured fields.
- The index status remains "Processing" or "Not Ready" for an extended period, indicated by the knowledge base being inaccessible to applications. This may be due to the
PARSE_FILE_TIMEOUT_SECONDSparameter being set too short, causing large or complex document formats to fail parsing.
How to Verify Configuration
- Upload a batch of typical deviation reports and CAPA records, then check if the knowledge base index status displays "Ready."
- Query specific deviation numbers, affected product batch numbers, or CAPA measures. Observe if the recalled document snippets accurately contain these critical entity information and compare them with the original documents.
- Test complex questions regarding root cause analysis or impact assessment. Evaluate the logical coherence and professionalism of the model's answers, confirming whether the answers are generated based on knowledge base content.
- Adjust the
Similarity Thresholdto observe changes in the quantity and quality of recall results, finding a reasonable range that balances recall rate and accuracy.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.