Citation and Traceability for Small Molecule Drug Pharmacovigilance

Pharmacovigilance data for small molecule drugs primarily originates from clinical trial reports, real-world evidence (RWE), post-market spontaneous

Data Characteristics

Pharmacovigilance data for small molecule drugs primarily originates from clinical trial reports, real-world evidence (RWE), post-market spontaneous adverse event reporting systems (e.g., FDA FAERS, EMA EudraVigilance), and medical literature. Data update frequencies vary. Clinical trial data typically releases after study completion, while post-market reporting systems continuously receive new data. Document structures are diverse, including structured Case Report Forms (CRF), semi-structured Individual Case Safety Report (ICSR) XML files, and unstructured medical literature abstracts and full texts. Key fields include drug name, active ingredient, dosage form, dose, indication, adverse event terminology (often MedDRA coded), onset time, outcome, and causality assessment. Units for dose are commonly milligrams (mg) or grams (g), and time is measured in days, months, or years.

Constraints on Citation and Traceability from Data Characteristics

The diversity of small molecule drug pharmacovigilance data imposes specific requirements on citation and traceability. Structured data, such as MedDRA codes, requires precise matching for accurate recall. Semi-structured and unstructured data necessitate more flexible text chunking and embedding strategies. Inconsistent data update frequencies demand that the knowledge base supports incremental updates and version management to ensure traced information is current. The broad range of document sources means integrating multiple external data sources and clearly identifying the original provenance of each information fragment. Specifically, key fields in adverse event reports, such as drugs, events, and causality, must trace back to a specific report ID or literature DOI when cited, supporting subsequent detailed review.

Configuration Settings

Configuration ItemRecommended ValueRationale
Chunk size (Chunk Length)500–800 charactersBalances the completeness of multi-field information in adverse event reports with the recall efficiency of individual chunks.
Recall count (Recall Count)Top 8Given the complexity of adverse event reports, increasing the recall count appropriately covers more potentially relevant information.
Similarity threshold (Similarity Threshold)0.75Addresses the need for precise matching of medical terminology, raising the threshold to filter out low-relevance results.
Rerank result count (Reranked Return Count)Top 5Further optimizes results through reranking after initial recall, providing the most relevant evidence.
PARSE_FILE_TIMEOUT_SECONDS300 secondsAccommodates parsing time for large adverse event database export files or complex medical literature.
maxContext2000 tokensEnsures the context window includes sufficient adverse event report details and related drug information.

Common Mistakes

  • Cited results contain outdated adverse event report information because the knowledge base does not regularly synchronize the latest data or lacks version control.
  • Traced source documents are inaccessible or content does not match, typically due to changes in file storage paths without corresponding updates in the knowledge base index.
  • Query results for specific drug-event pairs are too few, possibly due to an overly aggressive chunking strategy that splits critical information across different blocks.

How to Verify Configuration

  • For known updated drug adverse events, execute queries and verify that the publication date of the cited source aligns with the latest data source.
  • Randomly select document IDs or DOIs from cited sources and attempt to directly access the original data source, confirming content matches the fragment cited in the knowledge base.
  • Select multiple queries with clear associations between drugs and adverse event terms. Check if the number of returned citation entries meets expectations and evaluate their relevance.
  • Verify that the system accurately recalls corresponding structured data when processing queries containing standardized fields like MedDRA codes.

Note: The values provided are common starting points. Measure them against your own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.