Knowledge Base Retrieval and Recall for Pharmacovigilance Products

Pharmacovigilance data originates from clinical trial reports, real-world studies, adverse event reports (ADRs), drug inserts, medical literature, and

Data Characteristics in this Category

Pharmacovigilance data originates from clinical trial reports, real-world studies, adverse event reports (ADRs), drug inserts, medical literature, and regulatory guidelines. This data updates frequently, especially post-market adverse event reports, which exhibit continuous and fragmented characteristics. Document structures vary, including structured database records, semi-structured tabular data (e.g., adverse event lists in CSV or XLSX format), and extensive unstructured text (e.g., clinician notes, patient narratives, expert opinions). Data fields cover drug names, batches, indications, dosages, adverse reaction names, occurrence times, severity, prognosis, related diseases, and patient demographics. Units involve dosage (mg, g), frequency (times/day), and time (hours, days, years).

Constraints Imposed by these Characteristics on "Knowledge Base Retrieval and Recall"

The diversity and high update frequency of pharmacovigilance data require the knowledge base to support efficient incremental updates and multi-format file processing. Medical terms and abbreviations in unstructured text demand precise identification, posing challenges for tokenization and embedding models. The fragmented nature of adverse event reports means a single report may not provide complete information; aggregating multiple relevant documents is necessary for comprehensive risk assessment, which impacts recall strategies. Additionally, drug inserts and regulatory documents contain extensive tabular data. Key information (e.g., drug interactions, contraindications) must be accurately extracted from these tables and linked with unstructured text, directly affecting the effectiveness of knowledge base document chunking and metadata extraction. Sensitivity to units like time and dosage requires the retrieval system to understand and match numerical queries with specific units.

Configuration Settings

Configuration ItemRecommended ValueRationale
Chunk Length800–1200 charactersBalances the completeness of adverse event reports with the semantic coherence of long medical literature, preventing truncation of critical information.
Overlap Length100–200 charactersEnsures contextual continuity across paragraphs, improving retrieval recall, especially for text describing complex adverse reaction mechanisms.
Recall CountTop 10–15 itemsPharmacovigilance queries often require synthesizing information from multiple sources for judgment. Increasing recall count can cover more potentially relevant documents.
Similarity ThresholdCalibrate based on actual measurementsFor medical terms and abbreviations, evaluate and adjust using a test set to ensure high-relevance documents are recalled while filtering low-quality results.
Rerank CountTop 5 itemsWhile ensuring broad recall, the reranking model focuses on the most relevant, high-quality information, reducing subsequent analysis burden.
PARSE_FILE_TIMEOUT_SECONDS600 secondsAddresses long parsing times for large clinical trial reports or regulatory documents, preventing file processing failures due to timeouts.

Three Common Mistakes

  • New data is not retrieved after a knowledge base update. This happens because the text index was not rebuilt promptly or the indexing service was abnormal, leading to a { "message": "text index required for $text qu error.
  • Retrieval results contain many irrelevant or low-relevance documents. This occurs when Chunk Length is set too small, causing semantic loss, or Similarity Threshold is set too low, failing to effectively filter noise.
  • Key fields in tabular data (e.g., adverse event lists in XLSX files) are not effectively retrieved after importing into the knowledge base. This may be because the file parser failed to correctly identify the table structure and extract key columns as metadata or text content.

How to Verify Correct Configuration

  • Select representative pharmacovigilance queries. Simulate real-world scenarios to check if recall results include all known relevant documents.
  • For adverse event queries related to specific drugs, assess if retrieval results provide sufficient information to support risk assessment (e.g., adverse reaction names, incidence rates, related factors). Compare with manual query results to confirm coverage.
  • Upload files containing complex tabular structures. Check if the knowledge base document chunks correctly extract key information from tables and allow retrieval based on table content.

Note: The values provided are common starting points. Measure against your own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.