Knowledge Base Retrieval for Small Molecule Drug Pharmacovigilance

Small molecule drug pharmacovigilance data primarily originates from post-market surveillance reports, clinical trial reports, literature, drug

Data Characteristics

Small molecule drug pharmacovigilance data primarily originates from post-market surveillance reports, clinical trial reports, literature, drug labels, and global regulatory databases such as the FDA Adverse Event Reporting System (FAERS). This data updates frequently, especially post-market surveillance reports, exhibiting continuous and dynamic changes. Document structures are often semi-structured or unstructured text. For example, adverse event reports commonly include free-text descriptions and medical terminology codes (e.g., MedDRA). Typical fields include patient demographics, drug name, dosage, administration route, adverse event description, onset time, outcome, and causality assessment. Dosage units frequently involve mg, g, ml, and IU. Time units are precise to days or hours. There is also extensive mixing of medical jargon, drug brand names, and generic names.

Constraints on Knowledge Base Retrieval and Recall

The high update frequency of small molecule drug data requires an efficient incremental update mechanism for the knowledge base to ensure timely retrieval results. The prevalence of semi-structured and unstructured text means simple keyword matching struggles to capture deep semantic relationships, necessitating more sophisticated semantic understanding capabilities. Precise matching of medical and specialized terminology is crucial; general dictionaries may not cover all variations and abbreviations. Additionally, multi-dimensional information in adverse event descriptions (drug, event, time, dosage) demands that the knowledge base support complex queries and normalize different field units. Data format inconsistencies from multiple sources also challenge knowledge base preprocessing and index construction, requiring meticulous cleaning and standardization.

Configuration Guidelines

Configuration ItemSuggested ValueRationale
Chunk size (Segment Length)500–800 charactersEnsures each knowledge block contains sufficient context while avoiding excessive length that could dilute semantics or reduce recall efficiency.
Chunk Overlap Length (Segment Overlap Length)50–100 charactersMaintains context continuity, reducing semantic information loss due to segment truncation.
Recall count (Recall Count)8–12 itemsBalances coverage while avoiding recall of excessive irrelevant information, which would increase subsequent processing burden.
Similarity threshold (Similarity Threshold)0.75–0.85Balances recall accuracy and recall rate, reducing interference from irrelevant documents.
Rerank result count (Rerank Return Count)3–5 itemsFurther refines recall results, improving the quality of knowledge passed to the large language model.
Max Concurrent FilesCalibrate by measurementBalances system resource consumption and knowledge base update efficiency, especially when handling a large volume of report files.

Common Pitfalls

  • Knowledge base query results are empty, even if relevant information exists in the original documents. This happens when medical professional terms are not effectively expanded with synonyms or standardized, leading to a mismatch between query terms and document terms.
  • The knowledge base retrieves data, but the large language model's output does not reflect the knowledge base content. This can occur if Recall count (Recall Count) is set too low, or Similarity threshold (Similarity Threshold) is too high, resulting in fragmented or insufficient relevant knowledge to support the large language model's generation.
  • After uploading a compressed package containing many adverse event reports, some files fail to be ingested into the knowledge base. This may be due to PARSE_FILE_TIMEOUT_SECONDS being set too short, or individual file sizes exceeding the UPLOAD_FILE_MAX_SIZE limit, causing parsing timeouts or file rejection.

Configuration Validation

  • Query with keywords (e.g., specific drug names, adverse event terms). Check if the recall results include multiple relevant documents and if their Chunk size (Segment Length) and Chunk Overlap Length (Segment Overlap Length) meet expectations.
  • Use queries containing medical terminology abbreviations or aliases. Verify if the knowledge base successfully recalls documents with full terms, confirming the effectiveness of the synonym processing mechanism.
  • Upload a batch of simulated new adverse event reports. Observe if the knowledge base can immediately recall this new data through queries after updating, verifying the timeliness of the incremental update mechanism.
  • For specific adverse reactions, input complex queries involving multiple drugs, dosages, and onset times. Check if the recall results precisely match document segments containing this multi-dimensional information.

The values provided are common starting points and should be measured against specific samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.