Vector Models and Indexing for Small Molecule Drug Pharmacovigilance

Pharmacovigilance data for small molecule drugs originates from clinical trial reports, real-world evidence (RWE), post-market surveillance, medical

Data Characteristics

Pharmacovigilance data for small molecule drugs originates from clinical trial reports, real-world evidence (RWE), post-market surveillance, medical literature, and spontaneous patient reports. This data updates frequently, especially during a drug's initial market release, with new adverse event reports potentially generated weekly or even daily. Document structures typically include various types: structured reporting forms (e.g., ICH E2B format), semi-structured medical texts (e.g., patient records, diagnostic reports), and unstructured free-text descriptions. Key fields include drug name, active ingredient, dosage form, dose, administration route, adverse event (MedDRA coding), onset time, severity, outcome, relevant medical history, and concomitant medications. Units commonly used are milligrams (mg), grams (g), milliliters (mL) for dosage, and hours, days, weeks, months, years for time.

Constraints Imposed by Data Characteristics on Vector Models and Indexing

The high frequency of updates in small molecule drug adverse event reports demands that vector indexes possess efficient incremental update capabilities to ensure timely retrieval results. Data diversity, particularly the mix of structured data and free text, requires vector models to process and effectively integrate both types of information. The use of specialized terminology like MedDRA codes places higher demands on the model's semantic understanding to accurately capture medical concepts. Furthermore, adverse event descriptions often include time series and causal relationships, requiring the vectorization process to capture inter-event associations. Accurate matching and recall of key fields such as drug name, dose, and adverse events are fundamental to effective pharmacovigilance, making precise vector recall critical.

Configuration Guidelines

Configuration ItemRecommended ValueRationale
Chunk size (Chunk Length)300–500 characters (characters)Accommodates the typical length of factual descriptions in adverse event reports, avoiding information fragmentation.
Chunk Overlap Length (Overlap Length)50–80 characters (characters)Ensures context continuity and improves the accuracy of semantic connections across chunks.
Recall count (Recall Count)Top 10–20 entries (top 10–20 items)Balances recall rate with computational resource consumption, covering potentially relevant information.
Similarity threshold (Similarity Threshold)Calibrated by empirical measurementRequires evaluation using F1 score based on specific datasets and business needs.
embeddingModelbce-embeddingProvides good semantic understanding for Chinese medical texts.
Index Update Strategyincremental update, daily batch processingAddresses high-frequency data updates, ensuring data timeliness.

Common Pitfalls

  • Knowledge base search takes too long, manifesting as query requests with prolonged unresponsiveness or timeouts. This may be due to the chosen vector model having high computational resource requirements, while hardware configurations (e.g., CPU or GPU memory) are insufficient.
  • Retrieval results contain a large number of irrelevant or low-relevance documents. This may be due to a Similarity threshold (similarity threshold) set too low, leading to an overly broad recall scope.
  • Despite the upstream embedding service running normally, FastGPT reports "No Available channel" (no available channel). This usually means that the channel pointed to by ONEAPI_URL or FASTGPT_EMBEDDING_MODEL configuration items is not correctly mapped to an available embedding service.

Verification Steps

  • Using the FastGPT management interface, select the configured embeddingModel and perform vectorization tests on typical adverse event reports. Observe if the embedding field generates a reasonable numerical array.
  • Upload a batch of documents containing known adverse events. Use relevant drug names or adverse event symptoms to perform retrieval and check if the expected documents are included in the recall results, evaluating their ranking.
  • Monitor FastGPT operation logs to confirm that no connection errors or authentication failures related to the embedding service occur during knowledge base updates or queries.
  • For critical adverse event queries, perform multiple retrievals and adjust the Similarity threshold (similarity threshold) based on feedback from business experts regarding the relevance of retrieval results until satisfactory.

The values provided are common starting points and should be measured against the reader's own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.