Vector Models and Indexing for Antibody-Drug Conjugate (ADC) Products

Antibody-Drug Conjugate (ADC) data comes from various sources. These include public clinical trial databases, drug discovery literature, patent

Data Characteristics for this Category

Antibody-Drug Conjugate (ADC) data comes from various sources. These include public clinical trial databases, drug discovery literature, patent information, and internal pharmaceutical company compound libraries. Data update frequencies vary; clinical trial data typically updates when trials progress or results are published, while patent and literature data increases periodically with academic publications and approval cycles. ADC drug documentation structures are complex. They often contain target information, conjugation technology, toxin molecule structures, linker types, pharmacokinetic (PK) and pharmacodynamic (PD) data, and preclinical and clinical trial results. Key fields and units include molecular formula, CAS number, molar mass, and half-life. Examples of units are ng/mL (nanograms per milliliter) for blood concentration and nM (nanomolar) for binding affinity.

Constraints from these Characteristics on Vector Models and Indexing

ADC drugs have multimodal data, such as structural formulas, sequence information, and text descriptions. This requires vector models to effectively integrate different data types. High-dimensional molecular structure data needs specialized encoding strategies to capture functional groups and spatial conformation features. Irregular data updates, especially the phased disclosure of clinical trial data, mean that indexing must support incremental updates and version management to ensure information timeliness. Documents contain specialized terminology, abbreviations, and interdisciplinary concepts, such as immunology and organic chemistry. This demands higher semantic understanding capabilities from vector models. Additionally, numerical and time-series information in pharmacokinetic and pharmacodynamic data requires vector indexes to support multi-attribute filtering and range queries for precise retrieval of drug performance under specific conditions.

Configuration Settings

Configuration ItemRecommended ValueRationale
Chunk size (Segment Length)512 characters (512 characters)Balances semantic completeness with vector model processing efficiency. Prevents overly long segments from diluting key information.
Recall count (Recall Count)Top 10 entries (Top 10)Ensures sufficient relevant document snippets are retrieved in the initial recall phase, providing a basis for subsequent re-ranking.
Similarity threshold (Similarity Threshold)0.75-0.85Addresses the high number of specialized terms and strong semantic associations in the ADC domain, ensuring high relevance in results.
Rerank result count (Re-ranking Return Count)Top 3 entries (Top 3)Focuses on a few highly relevant, high-quality results, reducing redundant information.
PARSE_FILE_TIMEOUT_SECONDS600 seconds (600 seconds)Accommodates the parsing needs of PDF documents containing complex molecular structure diagrams and extensive tables.
embeddingModeltext-embedding-ada-002 or locally deployed bge-large-zhBalances semantic understanding depth with deployment flexibility, supporting processing of Chinese specialized literature.

Three Common Pitfalls

  • The knowledge base frequently displays "Indexing..." status and takes a long time to complete, leading to delayed query results. This can happen if the PARSE_FILE_TIMEOUT_SECONDS parameter is set too low. ADC research and development reports containing many images and complex tables may not be parsed within the allotted time.
  • Queries for specific molecular structures or pharmacokinetic data return insufficient relevance or irrelevant information. This occurs if the Similarity threshold (similarity threshold) is set too low, or if stop words or synonyms are not customized for ADC-specific terminology.
  • Calling the embedding model results in a 503 Service Unavailable error, with a message like "Current Group default For Model text-embedding" (for model text-embedding in current group default), causing vectorization to fail. This is typically due to One API routing configuration issues, or the backend embedding service being overloaded or experiencing connection timeouts.

How to Verify Configuration

  • Upload a typical ADC research and development report (including structural formulas, PK/PD curves). Check if the knowledge base successfully completes indexing without timeout or parsing errors.
  • Query for core concepts such as ADC targets, conjugation methods, and toxin types. Evaluate the quality and quantity of results returned based on Recall count (recall count) and Similarity threshold (similarity threshold).
  • Use queries containing precise fields like specific CAS numbers or molar masses. Confirm the system can return document snippets containing these fields through multi-attribute filtering or exact matching.
  • Check the logs of One API or the local embedding service. Ensure embedding model calls are successful and response times are acceptable during knowledge base updates or queries.

Note: The values provided are common starting points and should be measured against specific samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.