Data Characteristics in this Category
Pharmacovigilance data in medical record quality control primarily comes from Electronic Medical Record (EMR) systems and Hospital Information Systems (HIS). This includes patient visit records, physician orders, lab reports, imaging reports, surgical records, and nursing notes. These comprise both structured and unstructured text. Data is typically stored in standard formats like HL7 and CDA, alongside extensive free-text descriptions. Update frequency is high, with new data generated during each patient visit, examination, or medication event. Document structure is complex; a complete medical record can contain multiple sections and attachments. Fields involve generic drug names, brand names, dosages, administration methods, treatment durations, routes of administration, adverse reaction descriptions, occurrence times, and severity. Units must be precise, down to milligrams, milliliters, days, and times.
Constraints on Knowledge Base Retrieval and Recall
The diversity and complex structure of medical record data challenge knowledge base segmentation. This requires more refined text chunking strategies to prevent information loss or excessive redundancy. High-frequency updates necessitate efficient incremental update and index reconstruction mechanisms to ensure timely recall results. Pharmacovigilance-related descriptions are dispersed across various parts of the medical record. For example, adverse reactions might appear in the chief complaint, history of present illness, physician orders, or nursing notes. This makes single-keyword retrieval inefficient, demanding stronger semantic understanding capabilities. The precision required for fields and units means recall results must accurately match entities, avoiding confusion (e.g., distinguishing drug dosage from patient weight). Colloquial descriptions and medical abbreviations in free text also increase recall difficulty, requiring expansion mapping with medical dictionaries.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Chunk Length) | 800–1200 characters | Balances context completeness for long texts with retrieval efficiency for short texts, suitable for medical record descriptions. |
Chunk Overlap Length (Chunk Overlap Length) | 100–200 characters | Ensures contextual continuity at chunk boundaries, improving semantic coherence. |
Recall count (Number of Retrieved Chunks) | Top 5–8 | Considers the complexity of medical record data, increasing recall quantity to cover potentially relevant information. |
Similarity threshold (Similarity Threshold) | 0.75–0.85 | Balances recall rate and accuracy, avoiding interference from irrelevant information while not missing critical pharmacovigilance clues. |
Rerank result count (Number of Reranked Chunks) | Top 3 | Further prioritizes the most relevant results through a reranking model based on initial retrieval. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Addresses potentially long parsing times for large or complex medical record documents, preventing parsing failures due to timeouts. |
Common Pitfalls
- Duplicate indexes or an abnormal increase in chunk count often occur when minor document changes are not correctly identified as updates. Instead, they are treated as new documents uploaded repeatedly, or the chunking strategy deviates during incremental updates.
- Empty or inaccurate search results might stem from incorrect binding of knowledge base variable selection to the actual knowledge base name, or insufficient semantic matching between the query and knowledge base content.
- Inability to distinguish between different knowledge bases during retrieval, leading to mixed results. This typically happens when the retrieval scope is not explicitly specified using parameters like
Knowledge Base NameorKnowledge base ID(Knowledge Base ID), and the system defaults to searching all knowledge bases.
Verification Steps
- Upload representative medical record documents. Check if the
chunk countmeets expectations and ifchunk contentis semantically complete without obvious truncation. - Construct a series of queries for different types of adverse reaction descriptions. Observe the
Recall count(number of retrieved chunks) andsimilarity scoreto ensure relevant medical record snippets are recalled and to evaluate their ranking rationality. - Simulate actual user query scenarios. Observe whether the
recall resultscontain key drug information, dosage units, and adverse reaction symptoms from the medical record, and verify if this information accurately corresponds to the original text.
The values provided are common starting points. Measure against your own samples to determine optimal configurations.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.