Vector Model and Indexing for Antibody-Drug Conjugate (ADC) Pharmacovigilance

Antibody-Drug Conjugate (ADC) pharmacovigilance data originates from clinical trial reports, real-world studies, post-market adverse event reports

ADC Pharmacovigilance Data Characteristics

Antibody-Drug Conjugate (ADC) pharmacovigilance data originates from clinical trial reports, real-world studies, post-market adverse event reports (e.g., CIOMS I forms, MedWatch forms), academic literature, and drug prescribing information. Data update frequency is higher during clinical trials, with reporting cycles of weeks or months. Post-market data involves continuous, unscheduled reporting. Document structures typically include unstructured free text (e.g., adverse event details, patient history), semi-structured tabular data (e.g., drug dosage, administration route, adverse event codes like MedDRA terms), and structured patient demographic information. Key fields include drug generic name, batch number, adverse reaction description, onset time, severity, outcome, relevant laboratory indicators (e.g., liver and kidney function, complete blood count), concomitant medications, and medical history. Units involve dosage (mg/kg), time (days, hours), and laboratory results (U/L, g/L, mmol/L).

Constraints Imposed by Data Characteristics on Vector Models and Indexing

Unstructured free text descriptions in ADC pharmacovigilance data challenge vector models to accurately capture the subtle semantics of adverse events. The presence of semi-structured encodings like MedDRA terms requires vector models to understand natural language and to recognize and utilize contextual information from standardized medical terminology. The continuous and unscheduled nature of data updates means indexing must support incremental updates, avoiding frequent full rebuilds to maintain timeliness. Multi-source heterogeneous data leads to varying document lengths, from brief reports to detailed clinical records. This requires a segmentation strategy that ensures critical information is not truncated or diluted. Structured information such as adverse event severity and outcome needs to be reflected in the vector generation process, allowing high-value information to be prioritized during retrieval. Specific laboratory indicator values and units require the vectorization process to distinguish between numerical magnitudes and unit differences, avoiding confusion.

Configuration Guidelines

Configuration ItemRecommended ValueRationale
Chunk size800–1200 charactersBalances contextual completeness and vectorization efficiency, accommodating the length of descriptive text in ADC reports.
Chunk Overlap Length100–200 charactersEnsures semantic continuity across segments, preventing critical information from being split by segment boundaries.
embedding_modeltext-embedding-3-largeProvides higher semantic understanding and dimensionality, improving the discriminative power for medical terminology and complex descriptions.
Recall count8–15 entriesConsiders the complexity and diversity of ADC adverse events, increasing recall quantity to improve coverage.
Similarity threshold0.75–0.85Balances precise recall and appropriate generalization, avoiding false negatives or excessive recall of irrelevant information.
PARSE_FILE_TIMEOUT_SECONDS600 secondsAccommodates the potentially longer time required to parse large clinical trial reports or detailed adverse event description files.

Common Pitfalls

  • The knowledge base query results contain many irrelevant entries. This is due to a Similarity threshold set too low, leading to overly broad recall and a failure to effectively filter noise.
  • Uploading large PDF clinical trial reports results in a File Parsing Timeout error. This occurs when the PARSE_FILE_TIMEOUT_SECONDS value is insufficient to handle the file's complexity and size.
  • After updating some adverse event reports, retrieval results do not reflect the latest information. This happens when the knowledge base has not undergone incremental indexing or index reconstruction, causing the vector database to be out of sync with the original data.

Verification of Configuration

  • Select test data containing typical ADC adverse reaction descriptions. Perform queries and verify that the recalled results include all relevant key information.
  • Upload a mixed document containing MedDRA terms and free text descriptions. Observe whether it is correctly segmented and check if the vector representations of each segment can distinguish the semantics of medical terms.
  • Monitor the completion status of index reconstruction tasks after changing the embedding_model. Ensure all knowledge bases are updated to the new model.
  • Simulate high-concurrency query scenarios. Check system response times and compare them with baseline performance. Ensure that the Recall count and Similarity threshold settings do not cause performance bottlenecks.

The values provided are common starting points and should be measured against specific samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.