Knowledge Base Retrieval and Recall for Medical Record Quality Control in Pharmacovigilance

Pharmacovigilance data in medical record quality control primarily comes from Electronic Medical Record (EMR) systems and Hospital Information Systems

Data Characteristics in this Category

Pharmacovigilance data in medical record quality control primarily comes from Electronic Medical Record (EMR) systems and Hospital Information Systems (HIS). This includes patient visit records, physician orders, lab reports, imaging reports, surgical records, and nursing notes. These comprise both structured and unstructured text. Data is typically stored in standard formats like HL7 and CDA, alongside extensive free-text descriptions. Update frequency is high, with new data generated during each patient visit, examination, or medication event. Document structure is complex; a complete medical record can contain multiple sections and attachments. Fields involve generic drug names, brand names, dosages, administration methods, treatment durations, routes of administration, adverse reaction descriptions, occurrence times, and severity. Units must be precise, down to milligrams, milliliters, days, and times.

Constraints on Knowledge Base Retrieval and Recall

The diversity and complex structure of medical record data challenge knowledge base segmentation. This requires more refined text chunking strategies to prevent information loss or excessive redundancy. High-frequency updates necessitate efficient incremental update and index reconstruction mechanisms to ensure timely recall results. Pharmacovigilance-related descriptions are dispersed across various parts of the medical record. For example, adverse reactions might appear in the chief complaint, history of present illness, physician orders, or nursing notes. This makes single-keyword retrieval inefficient, demanding stronger semantic understanding capabilities. The precision required for fields and units means recall results must accurately match entities, avoiding confusion (e.g., distinguishing drug dosage from patient weight). Colloquial descriptions and medical abbreviations in free text also increase recall difficulty, requiring expansion mapping with medical dictionaries.

Configuration Guidelines

Configuration ItemRecommended ValueRationale
Chunk size (Chunk Length)800–1200 charactersBalances context completeness for long texts with retrieval efficiency for short texts, suitable for medical record descriptions.
Chunk Overlap Length (Chunk Overlap Length)100–200 charactersEnsures contextual continuity at chunk boundaries, improving semantic coherence.
Recall count (Number of Retrieved Chunks)Top 5–8Considers the complexity of medical record data, increasing recall quantity to cover potentially relevant information.
Similarity threshold (Similarity Threshold)0.75–0.85Balances recall rate and accuracy, avoiding interference from irrelevant information while not missing critical pharmacovigilance clues.
Rerank result count (Number of Reranked Chunks)Top 3Further prioritizes the most relevant results through a reranking model based on initial retrieval.
PARSE_FILE_TIMEOUT_SECONDS600 secondsAddresses potentially long parsing times for large or complex medical record documents, preventing parsing failures due to timeouts.

Common Pitfalls

  • Duplicate indexes or an abnormal increase in chunk count often occur when minor document changes are not correctly identified as updates. Instead, they are treated as new documents uploaded repeatedly, or the chunking strategy deviates during incremental updates.
  • Empty or inaccurate search results might stem from incorrect binding of knowledge base variable selection to the actual knowledge base name, or insufficient semantic matching between the query and knowledge base content.
  • Inability to distinguish between different knowledge bases during retrieval, leading to mixed results. This typically happens when the retrieval scope is not explicitly specified using parameters like Knowledge Base Name or Knowledge base ID (Knowledge Base ID), and the system defaults to searching all knowledge bases.

Verification Steps

  • Upload representative medical record documents. Check if the chunk count meets expectations and if chunk content is semantically complete without obvious truncation.
  • Construct a series of queries for different types of adverse reaction descriptions. Observe the Recall count (number of retrieved chunks) and similarity score to ensure relevant medical record snippets are recalled and to evaluate their ranking rationality.
  • Simulate actual user query scenarios. Observe whether the recall results contain key drug information, dosage units, and adverse reaction symptoms from the medical record, and verify if this information accurately corresponds to the original text.

The values provided are common starting points. Measure against your own samples to determine optimal configurations.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.