ADC Pharmacovigilance Data Characteristics
Antibody-Drug Conjugate (ADC) pharmacovigilance data originates from clinical trial reports, real-world studies, post-market adverse event reports (e.g., CIOMS I forms, MedWatch forms), academic literature, and drug prescribing information. Data update frequency is higher during clinical trials, with reporting cycles of weeks or months. Post-market data involves continuous, unscheduled reporting. Document structures typically include unstructured free text (e.g., adverse event details, patient history), semi-structured tabular data (e.g., drug dosage, administration route, adverse event codes like MedDRA terms), and structured patient demographic information. Key fields include drug generic name, batch number, adverse reaction description, onset time, severity, outcome, relevant laboratory indicators (e.g., liver and kidney function, complete blood count), concomitant medications, and medical history. Units involve dosage (mg/kg), time (days, hours), and laboratory results (U/L, g/L, mmol/L).
Constraints Imposed by Data Characteristics on Vector Models and Indexing
Unstructured free text descriptions in ADC pharmacovigilance data challenge vector models to accurately capture the subtle semantics of adverse events. The presence of semi-structured encodings like MedDRA terms requires vector models to understand natural language and to recognize and utilize contextual information from standardized medical terminology. The continuous and unscheduled nature of data updates means indexing must support incremental updates, avoiding frequent full rebuilds to maintain timeliness. Multi-source heterogeneous data leads to varying document lengths, from brief reports to detailed clinical records. This requires a segmentation strategy that ensures critical information is not truncated or diluted. Structured information such as adverse event severity and outcome needs to be reflected in the vector generation process, allowing high-value information to be prioritized during retrieval. Specific laboratory indicator values and units require the vectorization process to distinguish between numerical magnitudes and unit differences, avoiding confusion.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size | 800–1200 characters | Balances contextual completeness and vectorization efficiency, accommodating the length of descriptive text in ADC reports. |
Chunk Overlap Length | 100–200 characters | Ensures semantic continuity across segments, preventing critical information from being split by segment boundaries. |
embedding_model | text-embedding-3-large | Provides higher semantic understanding and dimensionality, improving the discriminative power for medical terminology and complex descriptions. |
Recall count | 8–15 entries | Considers the complexity and diversity of ADC adverse events, increasing recall quantity to improve coverage. |
Similarity threshold | 0.75–0.85 | Balances precise recall and appropriate generalization, avoiding false negatives or excessive recall of irrelevant information. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Accommodates the potentially longer time required to parse large clinical trial reports or detailed adverse event description files. |
Common Pitfalls
- The knowledge base query results contain many irrelevant entries. This is due to a
Similarity thresholdset too low, leading to overly broad recall and a failure to effectively filter noise. - Uploading large PDF clinical trial reports results in a
File Parsing Timeouterror. This occurs when thePARSE_FILE_TIMEOUT_SECONDSvalue is insufficient to handle the file's complexity and size. - After updating some adverse event reports, retrieval results do not reflect the latest information. This happens when the knowledge base has not undergone incremental indexing or index reconstruction, causing the vector database to be out of sync with the original data.
Verification of Configuration
- Select test data containing typical ADC adverse reaction descriptions. Perform queries and verify that the recalled results include all relevant key information.
- Upload a mixed document containing MedDRA terms and free text descriptions. Observe whether it is correctly segmented and check if the vector representations of each segment can distinguish the semantics of medical terms.
- Monitor the completion status of index reconstruction tasks after changing the
embedding_model. Ensure all knowledge bases are updated to the new model. - Simulate high-concurrency query scenarios. Check system response times and compare them with baseline performance. Ensure that the
Recall countandSimilarity thresholdsettings do not cause performance bottlenecks.
The values provided are common starting points and should be measured against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.