Antibody-Drug Conjugates (ADC) Pharmacovigilance: Citation and Traceability

Antibody-Drug Conjugate (ADC) pharmacovigilance data primarily originates from clinical trial reports, real-world evidence (RWE) studies, regulatory

ADC Data Characteristics

Antibody-Drug Conjugate (ADC) pharmacovigilance data primarily originates from clinical trial reports, real-world evidence (RWE) studies, regulatory agency reporting systems (e.g., FDA Adverse Event Reporting System, FAERS), and academic literature. Data update frequencies vary; clinical trial data typically releases in batches after trials conclude, RWE data may update continuously, and regulatory system data is continuously entered. Document structures are diverse, including structured database records, unstructured clinical study report PDFs, and medical journal articles. Key fields include patient demographics, ADC drug name, dosage, administration route, adverse event (AE) terminology (typically MedDRA codes), onset time, severity, outcome, and ADC drug-relatedness assessment. Dosage units commonly use milligrams per kilogram (mg/kg), and time units are days or weeks.

Constraints from Data Characteristics on Citation and Traceability

The diversity of ADC pharmacovigilance data imposes specific requirements on citation and traceability. First, multimodal data sources mean the knowledge base needs robustness in handling PDFs, structured text, and unstructured text to ensure complete information extraction. Second, the specialized and standardized nature of AE terminology (MedDRA) requires accurate handling of these medical proper nouns during tokenization and entity recognition to avoid semantic loss during segmentation. Inconsistent update frequencies, especially the continuous updates of regulatory databases, demand that the knowledge base supports incremental updates and version management to ensure the timeliness of cited information. Finally, the complex mechanism of action of ADC drugs can lead to lengthy and highly related adverse event descriptions, which challenges text segmentation granularity control; overly fine-grained segmentation may lose context, while overly coarse-grained segmentation affects retrieval accuracy.

Configuration Settings

Configuration ItemRecommended ValueRationale for Recommendation
Chunk size (Segment Length)500–800 charactersBalances the completeness of adverse event description context with retrieval efficiency, avoiding irrelevant information in overly long segments.
Chunk Overlap Length (Segment Overlap Length)100 charactersEnsures critical information spanning across segments is effectively captured during retrieval, especially for long sentences or complex causal relationships.
Recall count (Recall Count)Top 8–12 entriesConsidering the complexity and diversity of ADC adverse reactions, increasing the recall count improves the coverage of relevant information.
Similarity threshold (Similarity Threshold)0.75–0.85Addresses the need for precise matching of medical terminology, setting a higher threshold to filter for highly relevant knowledge snippets.
maxContext3000–4000 tokensEnsures large language models can process sufficiently long contexts for reasoning and tracing complex adverse events.
PARSE_FILE_TIMEOUT_SECONDS600 secondsAccommodates longer parsing times for clinical trial report PDFs and academic papers, preventing parsing failures due to timeouts.

Common Pitfalls

  • Knowledge base Q&A fails to generate question-answer pairs, instead directly inserting the original text: This typically results from improper text segmentation strategies, such as overly long segment lengths or segmentation logic failing to identify valid Q&A structures, making it difficult for the model to extract independent questions and answers.
  • The workbench knowledge base only retrieves one document per simple query: This might relate to a low Recall count (Recall Count) or Rerank result count (Reranked Return Count) configuration, causing the system to return only the most relevant snippet from a single document after initial retrieval or reranking.
  • Testing after model integration shows an error, but it runs in the citation: This could be due to discrepancies between model configuration parameter validation logic (e.g., API key, model name) in the test interface and the actual runtime invocation logic, or unstable network connectivity in the test environment.

Verification Steps

  • Submit a query containing complex adverse event descriptions and check if the returned citations include relevant snippets from multiple documents.
  • Ask a question about a known rare adverse event for a specific ADC drug and confirm if the knowledge base can accurately cite relevant research reports or regulatory database records.
  • Upload a new clinical trial report PDF, observe its parsing progress and the number of generated knowledge snippets, ensuring PARSE_FILE_TIMEOUT_SECONDS is set appropriately and parsing succeeds.
  • In the Q&A results, for answers involving MedDRA-coded adverse events, verify if the citation source can trace back to the original text containing that code and confirm the code's accuracy.

The values provided are common starting points and should be measured against specific samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.