Vector Models and Indexing for Supplier Audit Pharmacovigilance

Supplier audit data in pharmacovigilance primarily comes from audit reports, Corrective and Preventive Action (CAPA) documents, supplier qualification

Data Characteristics

Supplier audit data in pharmacovigilance primarily comes from audit reports, Corrective and Preventive Action (CAPA) documents, supplier qualification certificates, Standard Operating Procedures (SOPs), and historical adverse event records. These documents are typically in PDF, Word, or scanned image formats. Data update frequency is relatively low, occurring mainly after periodic audits or significant changes. Audit reports are highly structured, containing fields for audit findings, risk levels, and recommended actions. CAPA documents focus on problem descriptions, root cause analysis, implementation plans, and verification results. Fields may include drug batch numbers, supplier codes, audit dates, non-conformance numbers, and risk scores. Units are often dates, text descriptions, or enumerated values.

Constraints from Data Characteristics on Vector Models and Indexing

The low update frequency of supplier audit documents means less pressure on incremental updates after initial index construction. This allows more resources to be allocated to generating high-quality initial vectors. Documents contain significant structured information and key entities, such as supplier names, drug batches, and specific non-conformance descriptions. Vectorization must effectively encode this information. The presence of scanned documents requires high-quality Optical Character Recognition (OCR) before vectorization; otherwise, information loss will impact recall accuracy. Audit reports and CAPA documents are often lengthy, necessitating a reasonable segmentation strategy. This prevents single text blocks from being too large, diluting key information, or too small, leading to insufficient context. Fields like risk scores and non-conformance numbers, while numerical, have high semantic importance. The vector model must capture their association with text descriptions.

Configuration Recommendations

Configuration ItemRecommended ValueRationale
embeddingModelmultimodal-embedding-v1Handles multimodal information, such as charts in audit reports or scanned documents.
Chunk size (Segment Length)800–1200 charactersBalances contextual completeness with key information density.
Chunk Overlap Length (Segment Overlap Length)100 charactersEnsures continuity of context across segment boundaries.
PARSE_FILE_TIMEOUT_SECONDS600 secondsAccommodates OCR processing time for large audit reports and scanned documents.
Recall count (Recall Count)Top 10Ensures an adequate number of potentially relevant document blocks are recalled initially.
Similarity threshold (Similarity Threshold)Calibrated by actual measurementRequires adjustment based on actual query scenarios and data distribution.

Common Pitfalls

  • Document indexing fails to complete for an extended period. This often results from excessively large files or numerous images causing OCR processing timeouts, or PARSE_FILE_TIMEOUT_SECONDS being set too low.
  • Key entity information is missing from retrieval results. This typically occurs because an inappropriate vector model was chosen, failing to effectively process structured or semi-structured fields in the documents.
  • Multiple queries for the same audit finding yield inconsistent results. This may be due to an unreasonable document segmentation strategy, causing semantically similar information to be split into different vector blocks.

Validation Steps

  • Upload a typical supplier audit report and a CAPA document. Verify that indexing completes normally without error messages.
  • Compare key information from the source documents (e.g., non-conformance numbers, risk levels). Use retrieval queries to confirm this information is accurately recalled.
  • Construct complex queries including supplier names, drug batches, or audit dates. Check if the relevance of the recall results meets expectations and evaluate the completeness of the recalled document blocks.

The values provided are common starting points and should be measured against the reader's own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.