Data Characteristics for this Category
Antibody-Drug Conjugate (ADC) registration dossiers have unique characteristics. Data sources are extensive, including clinical trial reports (Phases I-III), non-clinical study reports (pharmacology, toxicology, pharmacokinetics), manufacturing process validation documents, quality control standards and testing methods, stability study data, and regulatory agency guidelines and Q&A documents. These materials have varying update frequencies. Clinical data may be updated quarterly or annually as trials progress, while manufacturing processes and quality standards are relatively stable, revised only when significant changes occur. Document structures are complex, often organized in ICH CTD (Common Technical Document) format, comprising multiple modules and hierarchical levels. Fields and units are diverse, such as dosage units (mg/kg), efficacy indicators (ORR, PFS), safety event codes (MedDRA) in clinical reports, and specific indicators like antibody concentration (mg/mL), drug-antibody ratio (DAR), and conjugation rate in quality documents.
Constraints Imposed by these Characteristics on "Knowledge Base Retrieval and Recall"
The complexity and diversity of ADC data impose specific requirements on knowledge base retrieval and recall. Multi-level document structures necessitate support for deep semantic understanding and contextual association to avoid erroneous recalls caused by shallow matching. Frequently updated clinical data require the knowledge base to have an efficient incremental update mechanism to ensure the timeliness of retrieval results. Unique biomedical fields and units, such as DAR values and MedDRA codes, require the tokenizer and embedding model to accurately identify and process these proper nouns and numerical values, preventing information loss by treating them as ordinary words. Furthermore, registration dossiers often contain numerous tables and images. The knowledge base must effectively parse and utilize this non-textual information, incorporating it into the retrieval scope to enhance recall comprehensiveness.
Configuration Strategy
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size | 500–800 characters | Balances semantic completeness and recall efficiency, avoiding segments that are too small or too large in information content. |
Recall count | 8–12 entries | Controls the context length fed to the LLM while ensuring information coverage, reducing computational cost. |
Similarity threshold | Calibrate by actual measurement | Calibrates for ADC-specific terminology and numerical values, ensuring highly relevant recall and filtering low-quality results. |
Rerank result count | 3–5 entries | Further refines recall results, improving the quality and relevance of information ultimately presented to the user. |
PARSE_FILE_TIMEOUT_SECONDS | 300 seconds | Addresses potentially long parsing times for large clinical trial reports and process documents. |
maxContext | 32000 token | Ensures capacity for multiple recalled segments and their context, supporting complex question answering. |
Three Common Pitfalls
- Retrieval results do not match the query, but the semantic score is high: This occurs because the embedding model inadequately understands ADC domain-specific terminology, leading to vector representations that deviate from actual semantics.
- After a knowledge base update, some new materials cannot be retrieved: This happens when the incremental indexing mechanism is improperly configured or the update frequency is too low, failing to include new data in the index promptly.
- Answer generation takes too long, leading to a poor user experience: This is caused by
Recall countbeing set too high, orChunk sizebeing too small, resulting in the need to process a large amount of fragmented context.
How to Confirm Proper Configuration
- Conduct test queries on core ADC concepts and specific indicators to check the relevance and accuracy of recall results.
- Upload the latest clinical trial data and guidelines, then immediately perform retrieval to verify the recall effectiveness of new content.
- Simulate user queries in different network environments, record, and analyze the average response time from query to answer.
- Review retrieval logs to verify whether
Similarity thresholdandRerank result counteffectively filtered out low-relevance content.
Note: The values provided are common starting points and should be measured against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.