Knowledge Base Retrieval and Recall for Intelligent Triage in Pharmacovigilance

Data for intelligent triage in pharmacovigilance primarily originates from drug labels, medical literature, adverse event reports (CIOMS I forms)

Data Characteristics in This Domain

Data for intelligent triage in pharmacovigilance primarily originates from drug labels, medical literature, adverse event reports (CIOMS I forms), clinical guidelines, drug interaction databases, and pharmacological research reports. This data updates frequently. For example, drug labels may be revised due to post-market surveillance results, and adverse event reports are continuously generated. Document structures typically include both structured and unstructured information. Structured sections contain fields such as drug name, active ingredient, indications, contraindications, dosage and administration, adverse reaction lists, and drug interactions. Unstructured sections include adverse reaction descriptions and clinical observation records. Fields and units have strong medical specificity, such as dosage units (mg, g, IU), frequency (times/day, QD), and adverse reaction terminology (MedDRA codes).

Constraints on Knowledge Base Retrieval and Recall

The high update frequency of pharmacovigilance data requires the knowledge base to quickly synchronize the latest information. Failure to do so can lead to outdated triage information, impacting patient safety. Drug labels and medical literature are lengthy and contain extensive medical terminology and abbreviations, posing challenges for chunking granularity. Chunks that are too long or too short can reduce retrieval accuracy. Free-text descriptions in adverse event reports are highly diverse, requiring robust semantic understanding for effective recall of relevant cases. Furthermore, the high specificity of medical entities like drug names and adverse reaction terms demands precise matching from the retrieval system, avoiding confusion from synonyms or near-synonyms. Combining multimodal data (e.g., dosage information in tables) with text data for retrieval also increases recall complexity. Data sensitivity (patient privacy) necessitates careful attention to data anonymization and compliance during processing and retrieval.

Configuration Guidelines

Configuration ItemRecommended ValueRationale
Chunk Length800–1200 charactersAccommodates paragraph lengths in medical literature and drug labels, balancing contextual completeness and retrieval efficiency.
Chunk Overlap Length100 charactersEnsures contextual continuity and captures key information spanning across chunks.
Recall CountTop 5–8 itemsProvides sufficient but not overwhelming initial retrieval results, considering the rigor of pharmacovigilance.
Similarity ThresholdCalibrated by empirical testingAdjusts based on the precise matching requirements of medical terminology and semantic similarity; an initial value of 0.75 can be set.
Rerank Return CountTop 3 itemsFurther refines the most relevant information from the recall results, improving triage accuracy.
PARSE_FILE_TIMEOUT_SECONDS300 secondsHandles parsing time for large PDF drug labels or reports, preventing timeout interruptions.

Common Pitfalls

  • After a knowledge base update, the triage system fails to reflect the latest drug information or adverse event reports in a timely manner. This occurs because the knowledge base synchronization mechanism is not configured for real-time or near real-time updates, leading to data retrieval delays.
  • When a user queries specific drug adverse reactions, the system's recalled results are inaccurate or overly broad. This can happen due to improper chunking granularity, causing critical information to be fragmented or buried in irrelevant content.
  • When importing a large number of structured adverse event report files (e.g., in Excel format), the system processes slowly or even fails. This may be due to the UPLOAD_FILE_MAX_SIZE parameter being set too low, preventing the handling of large files, or issues with the parser's recognition of specific fields.

How to Verify Configuration

  • Upload the latest version of a drug label PDF file, check for a status_code of 200, and verify that the parsed content is complete, free of garbled characters, and that key structured information (e.g., adverse reaction lists) is correctly extracted.
  • Conduct multiple rounds of questioning about a specific drug and its rare adverse reactions. Observe if the document_id of the recalled results points to the expected document, and evaluate the relevance of the recalled items to the query. The relevance threshold should meet business requirements.
  • Simulate user inquiries, entering questions that include drug dosages and frequencies with units. Check if the system accurately identifies these units and provides corresponding medication advice or warning information, verifying field recognition accuracy.
  • Import an Excel format adverse event report containing a large number of rows (e.g., 10000 rows). Check the import process logs to ensure no out_of_memory or timeout errors, and verify that the number of imported data entries matches the source file.

The values provided are common starting points and should be measured against the reader's own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.