Model Integration and Configuration for Deviation and CAPA in Pharmacovigilance

Deviation and Corrective and Preventive Action (CAPA) data in pharmacovigilance originate from anomaly reports during production, quality audit

Data Characteristics

Deviation and Corrective and Preventive Action (CAPA) data in pharmacovigilance originate from anomaly reports during production, quality audit records, laboratory test results, and patient adverse event reports. Data update frequency varies, typically synchronizing with production batches, event occurrences, or investigation progress. Document structures are a mix of structured and unstructured data. Structured sections include fields like event type, occurrence time, involved product, responsible department, CAPA number, and status. Unstructured sections cover detailed event descriptions, investigation reports, root cause analyses, and CAPA plans and implementation records. Time units are usually precise to date and hour. Quantity units may include batches, items, milligrams, or international units, depending on the subject of the deviation or CAPA.

Constraints on Model Integration and Configuration

The mixed structured and unstructured nature of Deviation and CAPA data imposes specific requirements on model integration. Detailed event descriptions and investigation reports in unstructured text require strong text understanding capabilities, which dictates the choice of embedding and rerank models. The variable data update frequency means that model training and knowledge base synchronization strategies must balance real-time needs with resource consumption, avoiding frequent full updates. Key identifiers in documents, such as CAPA numbers and product batches, require precise matching during retrieval. This impacts knowledge base segmentation strategies and metadata extraction. Data involves quality and compliance, demanding high accuracy and traceability in retrieval results. Therefore, similarity threshold and recall count settings require careful consideration to minimize false positives and false negatives.

Configuration Settings

Configuration ItemRecommended ValueRationale
Chunk Size800–1200 charactersEnsures context completeness for unstructured text like CAPA reports and investigation details.
Overlap Size150–250 charactersEnhances semantic continuity between paragraphs, especially when describing complex events.
embeddingModeltext-embedding-ada-002 or higherHandles detailed event descriptions and root cause analyses, improving semantic representation accuracy.
rerankModelCalibrate based on actual measurementsImproves retrieval ranking for deviation and CAPA texts in specific business scenarios.
Similarity Threshold0.75–0.85Guarantees relevance of retrieval results, reduces interference from irrelevant information, and lowers false positives.
Recall Count8–12 itemsLimits the number of retrieval results while ensuring coverage, reducing subsequent processing burden.

Common Configuration Mistakes

  • Frequent No available channel errors during model calls. This typically indicates improper OneAPI routing configuration or that the required model is not included in the model group.
  • Some PDF deviation reports upload with empty content. This can occur if the PDF file's internal encoding or structure is complex, preventing the parser from extracting text correctly.
  • Retrieval results contain many irrelevant or duplicate CAPA entries. This may be due to a similarity threshold that is too low or unreasonable knowledge base segmentation, leading to the recall of many low-quality chunks.

Configuration Validation

  • Upload representative deviation and CAPA documents. Check if the segmented content in the knowledge base is complete and semantically coherent.
  • Use different types of query statements (e.g., event description, CAPA number, product batch). Check if the recall count and similarity of retrieval results meet expectations.
  • Simulate real-world problems. Verify if the model can accurately extract key information from the knowledge base and generate feasible CAPA suggestions.
  • Monitor API call logs for embedding and rerank models. Confirm the model version matches expectations and observe response times.

The values provided are common starting points. Measure them against your own samples for optimal results.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.