Data Characteristics
Antibody-Drug Conjugate (ADC) pharmacovigilance data originates from clinical trial reports, real-world studies, post-market surveillance, and adverse event reporting systems from regulatory bodies. Data updates frequently, especially during initial market release and critical clinical phases. Document structures vary, including structured Case Report Forms (CRFs), unstructured medical texts (e.g., physician handwritten notes, patient interview records), and semi-structured regulatory reports (e.g., MedDRA-coded adverse event reports). In addition to common demographic information and medication history, specific fields of interest include antibody drug targets, conjugation methods, linker types, payload drug types, dosage, and administration regimens. Units involve dosage (mg/kg), concentration (μg/mL), and time (days, weeks, months).
Constraints Imposed by These Characteristics on Knowledge Base Retrieval and Recall
The highly specialized and diverse nature of ADC pharmacovigilance data demands high precision in knowledge base retrieval. Clinical trial reports contain extensive specialized terminology, abbreviations, and complex drug interaction descriptions, requiring stronger semantic understanding. The prevalence of unstructured text makes traditional keyword matching ineffective for recall; semantic matching based on embedding models is necessary. Frequent data updates require the knowledge base to support efficient incremental indexing and real-time update mechanisms to ensure retrieval result timeliness. ADC-specific molecular structure information, such as targets, linkers, and payload drugs, must serve as critical metadata for filtering and sorting during retrieval to avoid numerous irrelevant results. Additionally, statistical information like adverse event severity and frequency needs effective integration during recall to support risk assessment.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale for Recommendation |
|---|---|---|
Chunk size | 500–800 characters | Balances semantic completeness and recall efficiency; avoids excessive truncation of key information. |
Chunk Overlap Length | 100–150 characters | Ensures contextual continuity; reduces semantic loss due to segment boundaries. |
Recall count | 10–15 entries | Covers a broader range of potentially relevant knowledge points; provides sufficient candidates for re-ranking. |
Similarity threshold | Calibrate by actual measurement | Ensures relevance of recalled results; prevents interference from low-quality or irrelevant information. |
Rerank result count | 3–5 entries | Selects the most relevant knowledge snippets; reduces model processing load; improves response speed. |
metadata_filter | {"drug_class": "ADC"} | Precisely filters based on ADC-specific attributes; narrows the retrieval scope. |
Common Pitfalls
- Retrieval results include numerous adverse event details for non-ADC drugs. This typically occurs because
metadata_filteris not effectively configured or field names are inconsistent. - AI responses contain statements contradicting knowledge base content. This usually happens when
Similarity thresholdis set too low, leading to the recall of low-relevance knowledge snippets, or when the model generates a response automatically if the knowledge base has no hit. - After a knowledge base update, retrieval results still show old data. This indicates that the knowledge base's incremental indexing or real-time update mechanism did not trigger correctly.
How to Verify Configuration
- Select typical ADC adverse event query scenarios. Verify that the
metadata_filterfield accurately screens for ADC-related data in the retrieval results. - Check system logs to confirm
Recall countandRerank result countalign with expected settings. Evaluate the semantic relevance of recalled content to determine an appropriateSimilarity threshold. - After adding or updating ADC-related data in the knowledge base, immediately execute a query to confirm that new knowledge points are recalled promptly.
- For critical information known to exist in the knowledge base but not recalled by queries, check if
Chunk sizeandChunk Overlap Lengthsettings are appropriate to prevent key information from being split or lost.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.