Knowledge Base Retrieval and Recall for Pharmacovigilance in Medical Affairs

Medical affairs departments primarily source pharmacovigilance data from post-market surveillance reports, clinical trial reports, real-world study

Data Characteristics in This Category

Medical affairs departments primarily source pharmacovigilance data from post-market surveillance reports, clinical trial reports, real-world study data, medical literature, regulatory safety information, and internal adverse event databases. This data updates frequently; post-market adverse event reports, in particular, are continuously logged. Document formats vary, including structured database records, semi-structured tables (e.g., CIOMS I forms), and unstructured text reports (e.g., medical observation notes, patient interview records, expert opinions). Data fields include patient demographics, drug information (batch, dosage, administration), adverse event descriptions (symptoms, signs, severity), event outcomes, relevant medical history, and concomitant medications. Units involve dosage (mg, g, IU), frequency (times/day, week), and time (day, month, year), often accompanied by non-standardized medical terminology and abbreviations.

Constraints Imposed by These Characteristics on Knowledge Base Retrieval and Recall

High update frequency requires an efficient incremental update mechanism for the knowledge base. This prevents the retrieval of outdated or inaccurate information. Diverse document structures challenge the knowledge base's ability to process heterogeneous data. It needs to balance precise matching of structured fields with semantic understanding of unstructured text. The complexity and non-standardization of medical terminology make traditional keyword searches ineffective. Stronger semantic understanding is necessary to handle synonyms, near-synonyms, and abbreviations. Adverse event descriptions often contain long texts, requiring robust chunking strategies and context preservation. The accuracy of fields and units is critical for retrieving key information like drug dosages and event timelines. The knowledge base must identify and process numerical information with specific units to support more refined filtering and matching.

Configuration Settings

Configuration ItemSuggested ValueRationale
Chunk Length500–800 charactersMedical texts have strong contextual relevance. Chunks that are too short lose context; chunks that are too long add noise.
Chunk Overlap Length100–150 charactersEnsures critical information spanning across chunks is not fragmented, improving recall completeness.
Recall CountTop 8–12Pharmacovigilance issues often require multi-faceted information. Increasing recall count improves coverage.
Similarity ThresholdCalibrate by measurementAdjust through testing with a dataset, considering the complexity of medical terminology, to balance recall and precision.
Rerank Return CountTop 3–5Reranking more accurately filters the most relevant results, improving final answer quality.
UPLOAD_FILE_MAX_SIZE500 MBAddresses the need to upload large clinical reports and literature, preventing upload failures due to file size.
PARSE_FILE_TIMEOUT_SECONDS600 secondsHandles the time-consuming parsing of medical text in complex PDFs or scanned documents, preventing timeouts.

Three Common Mistakes

  • Knowledge base query results are empty or irrelevant. This often occurs due to insufficient standardization of medical terminology, leading to semantic matching failures.
  • Response speed significantly slows down after connecting to the knowledge base, manifesting as an HTTP 504 Gateway Timeout. This may stem from an excessively high recall count setting or improperly configured file parsing timeout.
  • Uploading large medical literature or reports results in system errors like offset is out of bounds or request entity too large. This indicates that file size or chunking limits exceed default system configurations.

How to Confirm Proper Configuration

  • Select representative pharmacovigilance queries. Check if recall results include all expected relevant documents and evaluate their relevance threshold.
  • Upload and parse typical medical reports of different types (PDF, DOCX, TXT) and sizes. Observe if file processing is smooth, without timeouts or error messages.
  • For queries containing specific dosages and time units, verify that the recall results accurately extract and match this numerical information with units.

The values provided are common starting points and should be measured against specific samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.