Knowledge Base Retrieval and Recall for Antibody-Drug Conjugate (ADC) Registration Dossier Preparation

Antibody-Drug Conjugate (ADC) registration dossiers have unique characteristics. Data sources are extensive, including clinical trial reports (Phases

Data Characteristics for this Category

Antibody-Drug Conjugate (ADC) registration dossiers have unique characteristics. Data sources are extensive, including clinical trial reports (Phases I-III), non-clinical study reports (pharmacology, toxicology, pharmacokinetics), manufacturing process validation documents, quality control standards and testing methods, stability study data, and regulatory agency guidelines and Q&A documents. These materials have varying update frequencies. Clinical data may be updated quarterly or annually as trials progress, while manufacturing processes and quality standards are relatively stable, revised only when significant changes occur. Document structures are complex, often organized in ICH CTD (Common Technical Document) format, comprising multiple modules and hierarchical levels. Fields and units are diverse, such as dosage units (mg/kg), efficacy indicators (ORR, PFS), safety event codes (MedDRA) in clinical reports, and specific indicators like antibody concentration (mg/mL), drug-antibody ratio (DAR), and conjugation rate in quality documents.

Constraints Imposed by these Characteristics on "Knowledge Base Retrieval and Recall"

The complexity and diversity of ADC data impose specific requirements on knowledge base retrieval and recall. Multi-level document structures necessitate support for deep semantic understanding and contextual association to avoid erroneous recalls caused by shallow matching. Frequently updated clinical data require the knowledge base to have an efficient incremental update mechanism to ensure the timeliness of retrieval results. Unique biomedical fields and units, such as DAR values and MedDRA codes, require the tokenizer and embedding model to accurately identify and process these proper nouns and numerical values, preventing information loss by treating them as ordinary words. Furthermore, registration dossiers often contain numerous tables and images. The knowledge base must effectively parse and utilize this non-textual information, incorporating it into the retrieval scope to enhance recall comprehensiveness.

Configuration Strategy

Configuration ItemRecommended ValueRationale
Chunk size500–800 charactersBalances semantic completeness and recall efficiency, avoiding segments that are too small or too large in information content.
Recall count8–12 entriesControls the context length fed to the LLM while ensuring information coverage, reducing computational cost.
Similarity thresholdCalibrate by actual measurementCalibrates for ADC-specific terminology and numerical values, ensuring highly relevant recall and filtering low-quality results.
Rerank result count3–5 entriesFurther refines recall results, improving the quality and relevance of information ultimately presented to the user.
PARSE_FILE_TIMEOUT_SECONDS300 secondsAddresses potentially long parsing times for large clinical trial reports and process documents.
maxContext32000 tokenEnsures capacity for multiple recalled segments and their context, supporting complex question answering.

Three Common Pitfalls

  • Retrieval results do not match the query, but the semantic score is high: This occurs because the embedding model inadequately understands ADC domain-specific terminology, leading to vector representations that deviate from actual semantics.
  • After a knowledge base update, some new materials cannot be retrieved: This happens when the incremental indexing mechanism is improperly configured or the update frequency is too low, failing to include new data in the index promptly.
  • Answer generation takes too long, leading to a poor user experience: This is caused by Recall count being set too high, or Chunk size being too small, resulting in the need to process a large amount of fragmented context.

How to Confirm Proper Configuration

  • Conduct test queries on core ADC concepts and specific indicators to check the relevance and accuracy of recall results.
  • Upload the latest clinical trial data and guidelines, then immediately perform retrieval to verify the recall effectiveness of new content.
  • Simulate user queries in different network environments, record, and analyze the average response time from query to answer.
  • Review retrieval logs to verify whether Similarity threshold and Rerank result count effectively filtered out low-relevance content.

Note: The values provided are common starting points and should be measured against specific samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.