Data Characteristics for this Category
Antibody-Drug Conjugate (ADC) regulatory submission data involves multiple data types from biologics, small molecule drugs, and conjugation technologies. Data sources include preclinical study reports, clinical trial reports, CMC (Chemistry, Manufacturing, and Control) documents, pharmacology and toxicology reports, quality control standards, and regulatory compliance documents. These materials typically exist as PDFs, Word documents, and structured data tables (e.g., Excel). Document update frequency is high, especially during clinical trial phases, where data continuously accumulates and undergoes revision. The document structure is complex, containing extensive specialized terminology, abbreviations, figures, and tables. Key fields include antibody sequence, payload structure, drug-antibody ratio (DAR), batch number, analytical methods, stability data, and pharmacokinetic parameters. Units encompass molar concentration (nM), dosage (mg/kg), time (h), temperature (℃), with extremely high demands for precision and consistency.
Constraints Imposed by these Characteristics on Vector Models and Indexing
The complexity of ADC regulatory submission data places specific demands on vector models and indexing. First, multimodal data sources and complex document structures necessitate vector models capable of processing text, tables, and graphical information; a single text embedding model may not capture all semantics. Second, the dense specialized terminology and abbreviations require vector models with deep domain expertise to avoid semantic drift or misunderstanding. High update frequency demands an indexing system that supports efficient incremental updates and version management, ensuring retrieved information is always current and accurate. Furthermore, the high requirements for precision and consistency mean similarity search threshold settings must be particularly fine-tuned to avoid recalling irrelevant or semantically ambiguous results, which could impact compliance judgments. Finally, the ability to identify and extract key fields and units challenges vectorization and query filtering strategies, requiring metadata filtering to effectively narrow the search scope.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size | 800–1200 characters | Balances contextual completeness with vector model processing capability, preventing long texts from diluting key information. |
Chunk Overlap Length | 100–200 characters | Ensures semantic coherence across segments, especially for complex descriptions and argumentative passages. |
Recall count | Top 20 entries | Increases initial recall scope to address potential weak matches caused by specialized terminology. |
Similarity threshold | Calibrated by actual measurements, range 0.75–0.85 | Addresses the precision requirements of ADC domain terminology, determined by testing with domain corpora. |
Rerank result count | Top 5 entries | Focuses on the most relevant content, reducing the engineer's screening burden and improving efficiency. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Handles parsing of large PDFs and complex tables, preventing timeouts that lead to indexing failures. |
Three Common Pitfalls
- Knowledge base status remains "indexing" for extended periods: This typically results from file parsing timeouts or vector model service connection interruptions, especially when processing large or complex ADC data files.
- Search result similarity values are abnormally high (e.g., 10000+): This indicates issues with vector model configuration or data normalization, causing similarity calculation results to exceed the expected range and preventing effective filtering.
- Retrieval results deviate significantly from expectations, with critical information missing: Potential causes include an unreasonable segmentation strategy, leading to important context being cut off, or insufficient domain knowledge in the vector model, preventing accurate understanding of specialized terminology in ADC documents.
How to Confirm Proper Configuration
- Upload typical ADC submission document samples. Check if the knowledge base indexing status completes normally and verify file parsing logs for any anomalies.
- Perform retrieval using query statements containing core ADC concepts. Check if the
similarityvalues of the returned results are within a reasonable range and can be effectively filtered by adjusting theSimilarity threshold(similarity threshold). - Conduct precise queries for key fields of specific ADC drugs (e.g., drug-antibody ratio, batch information). Verify if the recalled documents contain accurate information for these fields and check the relevance of results within the
Rerank result count(reranked return count).
Note: The values provided are common starting points. Measure against your own samples to determine optimal settings.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.