Pharmacovigilance: Knowledge Base Retrieval and Recall

Pharmacovigilance data originates from post-market surveillance reports, clinical trial data, medical literature, regulatory safety information, and

Data Characteristics in Pharmacovigilance

Pharmacovigilance data originates from post-market surveillance reports, clinical trial data, medical literature, regulatory safety information, and spontaneous patient reports. This data updates frequently, with some regulatory reports updating weekly or monthly. Document structures vary, including structured case report forms (e.g., CIOMS I forms), semi-structured medical journal articles, unstructured free-text reports, and scanned paper documents. Fields typically include patient demographics, drug names, dosages, administration methods, adverse event descriptions, event timestamps, outcomes, medical history, and concomitant medications. Units involve dosage (mg, g, IU), time (days, weeks, months), and frequency (times/day). The data also contains extensive medical terminology, abbreviations, and aliases.

Constraints on Knowledge Base Retrieval and Recall

The diverse nature of pharmacovigilance data requires the knowledge base to handle multiple file formats, especially recognizing scanned documents. High update frequency necessitates efficient incremental updates and version management. The mix of structured and unstructured data means simple keyword matching is insufficient; semantic understanding is also required. Extensive medical terminology, abbreviations, and aliases challenge the accuracy of text segmentation and word embedding models, potentially leading to incomplete recall or irrelevant information. The presence of time and dosage units requires the retrieval system to identify and process this numerical information, preventing misjudgments due to unit mismatches. Due to data sensitivity, knowledge base access control and data anonymization capabilities are critical for compliance.

Configuration Guidelines

Configuration ItemSuggested ValueRationale
Chunk size (Segment Length)800–1200 charactersPharmacovigilance reports and medical literature often contain detailed descriptions. Longer segments help maintain contextual integrity and prevent truncation of critical information.
Chunk Overlap Length (Segment Overlap Length)100–200 charactersEnsures sufficient overlap between adjacent segments to cover cross-sentence information that might be missed at segment boundaries, especially for adverse event descriptions.
Recall count (Recall Count)top 8–12 itemsPharmacovigilance queries often require more comprehensive background information. Increasing the recall count can improve the coverage of relevant information and reduce the risk of omissions.
Similarity threshold (Similarity Threshold)Calibrate based on actual measurements: 0.75–0.85The domain has many specialized terms, requiring high semantic similarity. An initially higher threshold can be set and then fine-tuned based on actual recall performance to balance precision and recall.
Rerank result count (Reranked Return Count)top 5 itemsAfter initial filtering, a reranking model further optimizes the order of recalled results, placing the most relevant items at the forefront to enhance user experience.
PARSE_FILE_TIMEOUT_SECONDS600 secondsOCR and parsing processes for complex documents like scanned PDFs can be time-consuming, requiring a longer timeout to prevent parsing failures.

Common Mistakes

  • The conversation interface call fails to return the referenced knowledge base ID because the interface configuration does not explicitly request the quote field or the corresponding parameter is not enabled.
  • After uploading scanned PDF e-books, text content is not recognized. Retrieval results are empty or irrelevant because the knowledge base's OCR service is not enabled or incorrectly configured.
  • During retrieval in the business system, many irrelevant or duplicate results appear. This is due to an inappropriate knowledge base indexing strategy, such as too small a Chunk size (Segment Length) or too low a Similarity threshold (Similarity Threshold).

Verification Steps

  • Upload multiple scanned PDFs containing typical adverse event descriptions. Check if text content is accurately identified and extracted, and attempt to retrieve key medical terms within them.
  • For a report known to contain a specific adverse event, construct a query and verify if the recalled results include the core information from that report. Check if the quote field correctly points to the original document.
  • Use queries containing medical term aliases or abbreviations. Verify if the knowledge base can recall relevant documents through semantic understanding (e.g., can querying "heart attack" (myocardial infarction) recall reports containing "Acute myocardial infarction" (acute myocardial infarction)?).
  • Simulate high-concurrency retrieval requests to the knowledge base. Monitor system response times and resource utilization to ensure stable performance under actual business loads.

The values provided are common starting points and should be measured against the reader's own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.