Vector Models and Indexing for Attenuated and Inactivated Vaccine Pharmacovigilance

Attenuated and inactivated vaccine pharmacovigilance data originates from clinical trial reports, real-world evidence (RWE) studies, post-market

Data Characteristics

Attenuated and inactivated vaccine pharmacovigilance data originates from clinical trial reports, real-world evidence (RWE) studies, post-market surveillance reports, spontaneous adverse event reporting systems (e.g., VAERS, EudraVigilance), and relevant medical literature. Data updates frequently. Adverse event reports increase significantly during large-scale vaccination campaigns. Document structures vary, including structured report forms, semi-structured clinical study summaries, and unstructured medical notes and patient descriptions. Fields cover basic patient information, vaccination history (batch number, vaccination date), adverse event onset time, symptom descriptions (e.g., fever, redness, allergic reaction), severity, outcome, and concomitant medications. Symptom descriptions often combine medical terminology and natural language. Units include dosage (μg), time (hours, days), and temperature (℃).

Constraints Imposed by These Characteristics on Vector Models and Indexing

The multi-source and heterogeneous nature of attenuated and inactivated vaccine pharmacovigilance data requires vector models with strong semantic understanding to extract key information from different text formats. High-frequency data streams demand real-time indexing and incremental update mechanisms to avoid frequent full rebuilds. The mix of structured and unstructured document features means a single text vectorization method may be insufficient; entity recognition and relation extraction techniques should be considered. The combination of medical terminology and natural language in symptom descriptions complicates vectorization, potentially causing synonyms or near-synonyms to be too far apart in the embedding space, affecting retrieval accuracy. Additionally, precise matching requirements for critical fields like batch numbers and vaccination dates necessitate vector indexes that support structured data filtering and hybrid retrieval.

Configuration Guidelines

Configuration ItemRecommended ValueRationale
Chunk size (Segment Length)512 characters (characters)Balances contextual completeness with vectorization efficiency, preventing long texts from diluting key information.
Recall count (Recall Count)Top 10 entries (top 10)Ensures sufficient potentially relevant document segments are covered in the initial recall phase, increasing the effective range for subsequent re-ranking.
Similarity threshold (Similarity Threshold)Calibrate based on actual measurementsVaccine adverse event semantics are sensitive; validation with actual data is needed to balance recall and accuracy.
Rerank result count (Re-ranking Return Count)Top 3 entries (top 3)After refinement by the re-ranking model, the top few results usually satisfy most query needs.
maxContext4000 tokenAdapts to common large language model context windows, ensuring retrieval results can be fully sent.
embed_model_nametext-embedding-v2 or bge-large-zhBalances semantic understanding capabilities with vector representation effectiveness for Chinese medical texts.

Three Common Mistakes

  • Key adverse event reports are missing from query results because document segments are too long. This dilutes critical information within a single report, making accurate matching difficult after vectorization.
  • Retrieved vaccine batch information is inaccurate. This may happen because the vector model does not weight or specially process specific fields when handling structured data, weakening their importance in the embedding space.
  • The system responds slowly to queries because index rebuilding frequency is too high, or incremental update mechanisms are not optimized, causing index operations and query operations to compete for resources.

How to Verify Configuration

  • Select a batch of queries with known specific adverse events. Verify if the recall results include the expected original report document IDs.
  • For adverse event queries related to different vaccine batches, check the accuracy of batch information in the returned results, ensuring consistency with query conditions.
  • Monitor index update and query response times. Ensure the system maintains acceptable performance during high-frequency data updates.
  • Perform manual evaluation to assess whether the retrieved document segments contain the core semantics of the query and if there is redundant information.

The values provided are common starting points. Measure them against your own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.