ADC Product Data Characteristics
Antibody-Drug Conjugate (ADC) product data originates from clinical study reports, patent literature, academic papers, drug inserts, regulatory approval documents, and internal R&D documentation. Data updates are relatively stable, typically aligning with clinical trial progress, new drug launches, patent expirations, or regulatory changes. Document structures are complex, containing extensive specialized terminology, chemical structures, biological activity data, pharmacokinetic parameters, toxicology information, and clinical efficacy data. Key fields include specific antibody targets, conjugation methods, toxin molecule types, drug-antibody ratio (DAR), indications, adverse reactions, dosage, and administration. Units involve unique biomedical measurements such as molar concentration (nM), dosage (mg/kg), time (h/day), and response rate (%).
Constraints on Knowledge Base Retrieval and Recall
The highly specialized and complex nature of ADC product data imposes multiple constraints on knowledge base retrieval and recall. The extensive specialized terminology demands strong semantic understanding; simple keyword matching often leads to missed or incorrect retrievals. Chemical structures and biological activity data, often presented as images or tables, are difficult for traditional text vectorization to process effectively, requiring multimodal recognition techniques. Although document update frequency is stable, each update may involve significant revisions of critical information, necessitating robust incremental updates and version management for the knowledge base. Furthermore, data from different sources may have inconsistent formats (e.g., PDF for clinical reports, XML for patents, Word for inserts), requiring powerful heterogeneous data processing capabilities. The specificity of fields necessitates customized entity recognition and relation extraction models to ensure retrieval precision and minimize irrelevant information.
Configuration Guidelines
| Configuration Item | Recommended Approach | Rationale |
|---|---|---|
Chunk size | 800–1200 characters | Balances semantic completeness and recall efficiency. Avoids diluting key information in long texts and losing context in short texts. |
Overlap Length | 100–200 characters | Ensures context continuity, handles critical information spanning paragraphs, and reduces semantic fragmentation from splitting. |
Recall count | Top 8–12 entries | ADC consultations often require multi-faceted information. Increasing recall count covers potential relevance. |
Similarity threshold | Calibrate by empirical testing | Requires multiple tests with the specific embedding model and dataset to balance precision and recall. |
Rerank result count | Top 5 entries | After re-ranking model optimization, ensures the most relevant core information is displayed first, improving user efficiency. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | ADC documents are complex and parsing is time-consuming. Increasing the timeout prevents data loss due to parsing interruptions. |
Common Pitfalls
- Phenomenon: Retrieval results include numerous biological agents or small molecule drug information unrelated to ADC drugs. Reason: The semantic embedding model fails to effectively distinguish specialized ADC terminology and context, leading to similarity calculation deviations.
- Phenomenon: When a user queries a specific side effect of an ADC drug, the system fails to recall relevant content or recalls incomplete information. Reason: During data import, the knowledge base failed to effectively extract and structure key fields like side effects or adverse reactions from unstructured text, or the segmentation strategy truncated critical information.
- Phenomenon: After a knowledge base update, some queries for the latest clinical data still return old version information. Reason: The knowledge base lacks an effective version management mechanism or has an incomplete incremental update strategy, leading to delayed index synchronization with new data.
Validation Steps
- For core ADC product queries, compare retrieval results with expert-provided reference answers to verify the accuracy and completeness of recalled content.
- Perform a series of queries containing specialized terminology and chemical structure descriptions to check if relevant document snippets are accurately identified and recalled.
- Simulate scenarios with newly published clinical data. After updating the knowledge base, test whether queries targeting the new data are correctly recalled.
- Use A/B testing to compare user satisfaction and query success rates under different
Recall countandSimilarity thresholdconfigurations.
Note: The values provided are common starting points. Measure them against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.