Knowledge Base Retrieval and Recall for Telemedicine Pharmacovigilance

Telemedicine pharmacovigilance data originates from patient remote consultation records, electronic prescriptions, medication logs, smart wearable

Data Characteristics

Telemedicine pharmacovigilance data originates from patient remote consultation records, electronic prescriptions, medication logs, smart wearable device monitoring data, and patient self-reported adverse event systems. This data updates frequently, with some transmitted in real-time. Document structures are diverse, including unstructured doctor-patient dialogue text, semi-structured electronic prescriptions (drug names, dosages, frequencies), structured physiological indicators (heart rate, blood pressure), and adverse event reports (symptom descriptions, onset times, severity). Fields are highly specific, such as drug generic names, batch numbers, administration routes, adverse event codes (e.g., MedDRA), and order execution status. Units include mg, mL, times/day, mmHg, bpm, etc.

Constraints on Knowledge Base Retrieval and Recall

High-frequency data updates require the knowledge base to support rapid incremental synchronization to ensure retrieval result timeliness. Heterogeneous data sources necessitate complex data cleaning, standardization, and entity recognition during knowledge base construction to effectively link drug information and adverse event descriptions from different sources. The mix of unstructured text and structured data places higher demands on retrieval models, requiring support for both semantic retrieval and precise field matching. Specifically, the presence of specialized terminology like MedDRA codes demands that the knowledge base understands medical concept hierarchies to avoid recall omissions due to incomplete term matching. Real-time requirements dictate that retrieval latency must remain low to support immediate decision-making by doctors or AI assistants.

Configuration Guidelines

Configuration ItemSuggested ValueRationale
Chunk size500–800 charactersPatient descriptions or doctor's orders in telemedicine are typically short; this avoids loss of context from excessive segmentation.
Chunk Overlap Length100–150 charactersEnsures semantic continuity between segments, especially when critical information like drug names or dosages spans segments.
Recall count8–12 entriesBalances retrieval efficiency and coverage, particularly when initially screening for potential adverse event information.
Similarity threshold0.75–0.85Addresses the precision requirements for medical terminology, improving the relevance of recall results and reducing noise.
Rerank result count3–5 entriesIn pharmacovigilance, the most relevant, limited information needs to be prioritized for doctors or systems.
PARSE_FILE_TIMEOUT_SECONDS600 secondsHandles parsing time for large patient archives or historical medication record PDFs.

Common Pitfalls

  • After uploading files to the knowledge base, query results are incomplete or lack critical information. This may occur if the file parser inadequately supports specific electronic medical record formats or medication logs, leading to incorrect extraction of structured fields or semantic breaks during text segmentation.
  • When retrieving adverse event information, the system fails to recall seemingly related cases described using different medical terminology. This happens if the knowledge base index does not sufficiently leverage medical dictionaries or ontologies, lacking understanding and expansion for synonyms, hypernyms/hyponyms, or MedDRA codes.
  • When creating a knowledge base, the dropdown for selecting a text understanding model is empty. This can be due to the backend-configured language model or embedding model service not starting correctly or incorrect API key configuration, preventing the model list from loading.

Validation Steps

  • Upload typical patient consultation records and electronic prescription PDFs. Test retrieving adverse event information for specific drugs. Verify that recall results include all relevant details from the document and that key entities (e.g., drug names, dosages) are correctly identified.
  • Query adverse events for the same drug using different expressions (e.g., generic name, brand name, medical slang). Observe the coverage and accuracy of recall results, ensuring the knowledge base handles the diversity of medical terminology.
  • Simulate high-concurrency retrieval requests. Monitor system response times to confirm that retrieval latency meets business requirements for real-time interactive telemedicine scenarios.

The values provided are common starting points and should be measured against your own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.