Vector Model and Indexing for Antibody-Drug Conjugate (ADC) Regulatory Submission Preparation

Antibody-Drug Conjugate (ADC) regulatory submission data involves multiple data types from biologics, small molecule drugs, and conjugation

Data Characteristics for this Category

Antibody-Drug Conjugate (ADC) regulatory submission data involves multiple data types from biologics, small molecule drugs, and conjugation technologies. Data sources include preclinical study reports, clinical trial reports, CMC (Chemistry, Manufacturing, and Control) documents, pharmacology and toxicology reports, quality control standards, and regulatory compliance documents. These materials typically exist as PDFs, Word documents, and structured data tables (e.g., Excel). Document update frequency is high, especially during clinical trial phases, where data continuously accumulates and undergoes revision. The document structure is complex, containing extensive specialized terminology, abbreviations, figures, and tables. Key fields include antibody sequence, payload structure, drug-antibody ratio (DAR), batch number, analytical methods, stability data, and pharmacokinetic parameters. Units encompass molar concentration (nM), dosage (mg/kg), time (h), temperature (℃), with extremely high demands for precision and consistency.

Constraints Imposed by these Characteristics on Vector Models and Indexing

The complexity of ADC regulatory submission data places specific demands on vector models and indexing. First, multimodal data sources and complex document structures necessitate vector models capable of processing text, tables, and graphical information; a single text embedding model may not capture all semantics. Second, the dense specialized terminology and abbreviations require vector models with deep domain expertise to avoid semantic drift or misunderstanding. High update frequency demands an indexing system that supports efficient incremental updates and version management, ensuring retrieved information is always current and accurate. Furthermore, the high requirements for precision and consistency mean similarity search threshold settings must be particularly fine-tuned to avoid recalling irrelevant or semantically ambiguous results, which could impact compliance judgments. Finally, the ability to identify and extract key fields and units challenges vectorization and query filtering strategies, requiring metadata filtering to effectively narrow the search scope.

Configuration Settings

Configuration ItemRecommended ValueRationale
Chunk size800–1200 charactersBalances contextual completeness with vector model processing capability, preventing long texts from diluting key information.
Chunk Overlap Length100–200 charactersEnsures semantic coherence across segments, especially for complex descriptions and argumentative passages.
Recall countTop 20 entriesIncreases initial recall scope to address potential weak matches caused by specialized terminology.
Similarity thresholdCalibrated by actual measurements, range 0.75–0.85Addresses the precision requirements of ADC domain terminology, determined by testing with domain corpora.
Rerank result countTop 5 entriesFocuses on the most relevant content, reducing the engineer's screening burden and improving efficiency.
PARSE_FILE_TIMEOUT_SECONDS600 secondsHandles parsing of large PDFs and complex tables, preventing timeouts that lead to indexing failures.

Three Common Pitfalls

  • Knowledge base status remains "indexing" for extended periods: This typically results from file parsing timeouts or vector model service connection interruptions, especially when processing large or complex ADC data files.
  • Search result similarity values are abnormally high (e.g., 10000+): This indicates issues with vector model configuration or data normalization, causing similarity calculation results to exceed the expected range and preventing effective filtering.
  • Retrieval results deviate significantly from expectations, with critical information missing: Potential causes include an unreasonable segmentation strategy, leading to important context being cut off, or insufficient domain knowledge in the vector model, preventing accurate understanding of specialized terminology in ADC documents.

How to Confirm Proper Configuration

  • Upload typical ADC submission document samples. Check if the knowledge base indexing status completes normally and verify file parsing logs for any anomalies.
  • Perform retrieval using query statements containing core ADC concepts. Check if the similarity values of the returned results are within a reasonable range and can be effectively filtered by adjusting the Similarity threshold (similarity threshold).
  • Conduct precise queries for key fields of specific ADC drugs (e.g., drug-antibody ratio, batch information). Verify if the recalled documents contain accurate information for these fields and check the relevance of results within the Rerank result count (reranked return count).

Note: The values provided are common starting points. Measure against your own samples to determine optimal settings.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.