Knowledge Base Retrieval and Recall for mRNA Vaccine Pharmacovigilance

mRNA vaccine pharmacovigilance data originates from clinical trial reports, real-world evidence (RWE) data, global adverse drug reaction databases

Data Characteristics

mRNA vaccine pharmacovigilance data originates from clinical trial reports, real-world evidence (RWE) data, global adverse drug reaction databases (e.g., WHO VigiBase, FDA VAERS, EMA EudraVigilance), and medical literature. Data updates frequently, especially during initial market release or when new safety signals emerge. Documents are often unstructured or semi-structured text, including case reports, medical abstracts, and expert assessment reports, containing extensive free-text descriptions. Key fields include patient demographics (age, sex), vaccine batch number, vaccination date, adverse reaction onset date, adverse reaction description (MedDRA coding), severity, outcome, relevant laboratory test results, and concomitant medications. Adverse reaction descriptions typically include detailed clinical symptoms, signs, and diagnostic information, with units such as dose (μg), time (days, hours), and frequency (times/day).

Constraints on Knowledge Base Retrieval and Recall

Frequent updates require the knowledge base to support rapid incremental indexing, ensuring timely and accurate retrieval results. The large volume of unstructured text necessitates robust text parsing and semantic understanding capabilities to effectively extract adverse reaction association information. The presence of specialized terminology like MedDRA codes demands that the retrieval system understands medical terms and handles synonyms and hierarchical relationships to improve recall. The need for precise matching of key fields such as vaccine batch numbers and vaccination dates imposes requirements on metadata filtering and structured information retrieval. Additionally, the long text characteristics in adverse reaction descriptions require a chunking strategy that effectively preserves contextual semantics, preventing truncation of critical information that could affect retrieval accuracy. Data source diversity also increases the complexity of data standardization and deduplication.

Configuration Guidelines

Configuration ItemRecommended ValueRationale
Chunk size500–800 charactersBalances the detail of adverse reaction descriptions with retrieval efficiency, avoiding excessive length that introduces noise or excessive brevity that loses context.
Recall countTop 10–20 entriesEnsures coverage of potentially relevant information, balancing recall rate with the computational cost of subsequent reranking.
Similarity threshold0.75–0.85Filters out low-relevance results, preventing the introduction of inaccurate or irrelevant clinical cases. Specific values require empirical validation.
Rerank result countTop 5 entriesSelects the most relevant cases from the retrieved results, focusing on core adverse reaction information.
MaxTokens2048Accommodates the input requirements of long adverse reaction reports, ensuring complete context is passed to the large language model.
UPLOAD_FILE_MAX_SIZE100 MBSupports the upload requirements for large clinical reports and research documents.

Common Pitfalls

  • Observation: Retrieval results contain numerous irrelevant or low-relevance adverse reaction reports. Reason: The Similarity threshold (similarity threshold) is set too low, failing to effectively filter out noisy data.
  • Observation: Knowledge base search tests return results correctly, but the reranking model is not used during actual large language model invocation. Reason: The reranking model configuration or dependencies in the knowledge base search test environment differ from those in the large language model invocation environment, causing the reranking function to be inactive in specific scenarios.
  • Observation: Inability to trace the specific source document or passage of adverse reaction information in AI-generated responses. Reason: The knowledge base chunking strategy does not carry sufficient metadata (e.g., document ID, page number), or the retrieval results do not pass this metadata to the large language model.

Verification Steps

  • Simulate queries with different adverse reaction symptom descriptions. Check if the retrieved results include known highly relevant clinical cases and verify if the number of recalled items is within the expected range.
  • Randomly select AI-generated responses about adverse reactions. Trace the quoted knowledge base text to verify the match between the quoted content and the response, and check if the specific document source can be located.
  • Conduct retrieval tests for medical terms known to have polysemy or synonyms. Confirm that the system correctly understands the semantics and recalls all relevant variant information. Assess if the Chunk size (chunk length) and Similarity threshold (similarity threshold) settings are appropriate.
  • Simulate data update scenarios. Observe if newly added or updated documents are indexed promptly and participate in retrieval, verifying the real-time update capability of the knowledge base.

Note: The values provided are common starting points. Measure against your own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.