Tool Calling and Plugins for Lead Optimization in Pharmacovigilance

Pharmacovigilance data in the lead optimization phase originates from preclinical studies, early clinical trials (e.g., Phase I, Phase II), and in

Data Characteristics

Pharmacovigilance data in the lead optimization phase originates from preclinical studies, early clinical trials (e.g., Phase I, Phase II), and in vitro/in vivo pharmacology and toxicology reports. This data updates frequently, potentially weekly or even daily, especially with rapid research progress. Document structures typically include detailed experimental protocols, raw data records, data analysis reports, and conclusions. Data fields cover compound structure information, administration routes, dosages, subject exposure, descriptions of observed adverse events, severity, onset time, duration, drug-relatedness assessments, and biomarker data. Units include concentration (e.g., nM, μM), dosage (e.g., mg/kg), time (e.g., h, day), and effect values (e.g., IC50, LD50). High-throughput data, such as genomics and proteomics, may also be present, typically in structured files (e.g., .csv, .tsv) or semi-structured documents (e.g., experiment logs in .pdf).

Constraints on Tool Calling and Plugins

The co-existence of highly structured and semi-structured lead optimization data demands flexible data parsing capabilities from tool calling and plugins. For example, processing raw experimental reports in .pdf format requires plugins to accurately extract key information, such as adverse event descriptions or specific biomarker values. High data update frequency means the knowledge base needs frequent incremental updates or re-indexing to ensure the AI Agent accesses the latest information. For numerical data involving various units, tool calls require unit conversion or standardization to prevent errors due to unit inconsistencies. For instance, dosage information for one adverse event might be recorded in mg/kg, while related data uses μM. High-throughput data, due to its large volume and complexity, places high demands on plugin data processing capabilities and performance, potentially requiring specialized pre-processing tools or interfaces for integration.

Configuration Settings

Configuration ItemSuggested ValueRationale
UPLOAD_FILE_MAX_SIZE500 MBProvides sufficient upload space, considering that high-throughput data files can be large.
maxContext3000 TokensEnsures the inclusion of key information from longer experimental reports or multiple related documents, preventing context truncation.
Chunk size (Segment Length)800–1200 characters (characters)Balances text block integrity and recall efficiency, suitable for paragraph structures in experimental reports.
Recall count (Recall Count)Top 8 entries (top 8 items)Given the potentially high correlation of information in the lead optimization phase, increasing the recall count improves relevance coverage.
Similarity threshold (Similarity Threshold)Calibrate by actual measurement (calibrate by empirical measurement)Requires testing against different data sources and query types to effectively filter noise and retrieve relevant results.
PARSE_FILE_TIMEOUT_SECONDS600 seconds (seconds)Requires a longer parsing time when processing large or complex .pdf documents.

Common Mistakes

  • Inconsistent dosage or concentration values returned after tool calls occur because original data units are not uniformly handled, leading to confusion during tool parsing.
  • After a knowledge base update, the Agent still references old experimental results. This happens when the incremental update strategy is not configured correctly, or indexing frequency is insufficient.
  • When processing .pdf experimental reports, some critical information (e.g., adverse event severity) is not extracted. This typically indicates that the plugin's parsing capability for non-standard formats or complex tables is insufficient.

Verification of Configuration

  • For typical adverse event queries, verify that experimental data cited in Agent responses precisely matches values in original reports.
  • After uploading new experimental reports, use query functions to confirm the Agent can immediately access and correctly utilize the latest data. Check the knowledge base update timestamp.
  • Test with .pdf files containing complex charts or irregular text layouts to confirm the plugin accurately extracts all expected field values.
  • Simulate dosage queries with different units (e.g., nM, μM, mg/kg) to verify the Agent's consistency when handling unit-converted or standardized values.

The values provided are common starting points. Measure them against specific samples to determine optimal settings.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.