Data Characteristics
Deviation and Corrective and Preventive Action (CAPA) data in pharmacovigilance originate from anomaly reports during production, quality audit records, laboratory test results, and patient adverse event reports. Data update frequency varies, typically synchronizing with production batches, event occurrences, or investigation progress. Document structures are a mix of structured and unstructured data. Structured sections include fields like event type, occurrence time, involved product, responsible department, CAPA number, and status. Unstructured sections cover detailed event descriptions, investigation reports, root cause analyses, and CAPA plans and implementation records. Time units are usually precise to date and hour. Quantity units may include batches, items, milligrams, or international units, depending on the subject of the deviation or CAPA.
Constraints on Model Integration and Configuration
The mixed structured and unstructured nature of Deviation and CAPA data imposes specific requirements on model integration. Detailed event descriptions and investigation reports in unstructured text require strong text understanding capabilities, which dictates the choice of embedding and rerank models. The variable data update frequency means that model training and knowledge base synchronization strategies must balance real-time needs with resource consumption, avoiding frequent full updates. Key identifiers in documents, such as CAPA numbers and product batches, require precise matching during retrieval. This impacts knowledge base segmentation strategies and metadata extraction. Data involves quality and compliance, demanding high accuracy and traceability in retrieval results. Therefore, similarity threshold and recall count settings require careful consideration to minimize false positives and false negatives.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk Size | 800–1200 characters | Ensures context completeness for unstructured text like CAPA reports and investigation details. |
Overlap Size | 150–250 characters | Enhances semantic continuity between paragraphs, especially when describing complex events. |
embeddingModel | text-embedding-ada-002 or higher | Handles detailed event descriptions and root cause analyses, improving semantic representation accuracy. |
rerankModel | Calibrate based on actual measurements | Improves retrieval ranking for deviation and CAPA texts in specific business scenarios. |
Similarity Threshold | 0.75–0.85 | Guarantees relevance of retrieval results, reduces interference from irrelevant information, and lowers false positives. |
Recall Count | 8–12 items | Limits the number of retrieval results while ensuring coverage, reducing subsequent processing burden. |
Common Configuration Mistakes
- Frequent
No available channelerrors during model calls. This typically indicates improper OneAPI routing configuration or that the required model is not included in the model group. - Some PDF deviation reports upload with empty content. This can occur if the PDF file's internal encoding or structure is complex, preventing the parser from extracting text correctly.
- Retrieval results contain many irrelevant or duplicate CAPA entries. This may be due to a
similarity thresholdthat is too low or unreasonable knowledge base segmentation, leading to the recall of many low-quality chunks.
Configuration Validation
- Upload representative deviation and CAPA documents. Check if the segmented content in the knowledge base is complete and semantically coherent.
- Use different types of query statements (e.g., event description, CAPA number, product batch). Check if the
recall countandsimilarityof retrieval results meet expectations. - Simulate real-world problems. Verify if the model can accurately extract key information from the knowledge base and generate feasible CAPA suggestions.
- Monitor API call logs for
embeddingandrerankmodels. Confirm the model version matches expectations and observe response times.
The values provided are common starting points. Measure them against your own samples for optimal results.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.