Knowledge Base Retrieval and Recall for High-Value Consumables Pharmacovigilance

High-value consumable pharmacovigilance data primarily originates from hospital electronic medical record systems, adverse event reporting platforms

Data Characteristics for This Category

High-value consumable pharmacovigilance data primarily originates from hospital electronic medical record systems, adverse event reporting platforms, manufacturer complaint records, and regulatory agency announcements. Data updates frequently, especially when new products launch or widespread adverse events occur. Document types are diverse, including structured adverse event report forms, unstructured handwritten notes from medical staff, patient interview records, device batch information, instruction manual revision histories, and clinical research reports. Beyond standard patient information, adverse event descriptions, and diagnostic results, fields also include device-specific batch numbers, serial numbers, manufacturing dates, expiration dates, implantation sites, departments of use, and operating physicians. Units involve dimensions (millimeters), weight (grams), time (hours, days), quantity (pieces), and temperature (Celsius).

Constraints on Knowledge Base Retrieval and Recall

The fragmented and diverse nature of high-value consumable data sources requires the knowledge base to effectively integrate structured and unstructured information and perform multimodal matching during retrieval. High-frequency data updates mean the knowledge base needs to support efficient incremental indexing and real-time update mechanisms to ensure the timeliness of retrieval results. The presence of unique identifiers like batch numbers and serial numbers demands that the retrieval system accurately match specific batch product information and differentiate subtle variations between products. The unstructured nature of handwritten medical notes and patient interviews places higher demands on the accuracy of semantic retrieval, requiring handling of colloquial expressions, typos, and mixed professional terminology. Additionally, diverse fields and units must be standardized during knowledge base construction to avoid retrieval ambiguity, especially concerning critical parameters like dosage and dimensions.

Configuration Settings

Configuration ItemRecommended ValueRationale
Chunk size (Chunk Length)800–1200 characters (characters)Balances contextual completeness with retrieval granularity, suitable for adverse event descriptions.
Overlap Length100 characters (characters)Ensures contextual coherence and prevents truncation of key information.
Recall count (Recall Count)Top 8 entries (top 8)Covers potentially relevant information while considering subsequent model processing efficiency.
Similarity threshold (Similarity Threshold)0.75Filters highly relevant knowledge blocks and reduces noise.
Embedding Modelbge-large-zh-v1.5Optimized for Chinese medical terminology, improving semantic understanding.
UPDATE_INTERVAL_SECONDS3600 seconds (seconds)Ensures the knowledge base can respond to new adverse event reports within one hour.

Common Pitfalls

  • Retrieval results do not match the query, showing generalized or irrelevant content. This occurs due to improper knowledge base chunking strategies, leading to truncated key information or insufficient context, which affects embedding vector quality.
  • The system returns only Chinese answers for English queries or English knowledge base documents. This happens when the response_language parameter is not set correctly during model inference, or the local model used inherently has a language bias.
  • Semantic retrieval scores are unusually high, but the returned content quality is poor. This indicates a domain mismatch between the embedding model and the knowledge base document content, leading to distorted similarity calculations.

Validation Steps

  • Create multiple question-answer pairs for typical high-value consumable adverse event cases. Perform retrieval and manually verify the relevance and completeness of the recalled results.
  • Upload documents containing information on new batches or versions of consumables. Observe whether relevant queries can accurately retrieve the latest data after the knowledge base updates.
  • Use queries containing specialized terminology, abbreviations, or colloquial descriptions. Evaluate the retrieval system's ability to understand non-standard language and compare it against manually annotated expected results.

The values provided are common starting points and should be measured against specific samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.