Knowledge Base Retrieval and Recall for Antibody-Drug Conjugate (ADC) Quality Documents

Antibody-Drug Conjugate (ADC) quality documents typically include batch reports, method validation files, stability study data, and raw and auxiliary

Data Characteristics

Antibody-Drug Conjugate (ADC) quality documents typically include batch reports, method validation files, stability study data, and raw and auxiliary material release records. These documents are primarily in PDF, Word, or Excel formats. They have complex internal structures, often containing charts, tables, and extensive technical terminology. Data update frequency relates to the product lifecycle; for example, batch reports are generated with each production batch, and stability data is continuously updated at predefined time points. Document fields include batch number, production date, expiration date, test item, test result, unit (e.g., ug/mL, %), and quality control limits. Multiple revisions are common.

Constraints on Knowledge Base Retrieval and Recall

The complex structure and specialized terminology of ADC quality documents challenge knowledge chunking, requiring care to avoid semantic fragmentation. Multiple revisions mean the knowledge base must effectively manage version information to ensure retrieval of the latest or specified version. Structured data in Excel tables may lose context during traditional text chunking, affecting retrieval accuracy. Frequently updated data requires the knowledge base to have an efficient incremental update mechanism to avoid reprocessing large amounts of unchanged content. Additionally, numerical information, such as specific units and quality control limits, requires the vectorization model to capture their intrinsic meaning for distinguishing subtle quality differences during retrieval.

Configuration Settings

Configuration ItemRecommended ValueRationale
Chunk size (Chunk Length)500–800 charactersBalances semantic completeness and recall efficiency, preventing overly long paragraphs from diluting key information.
Chunk overlap (Chunk Overlap)50 charactersEnsures contextual continuity and reduces loss of critical information due to chunking.
Recall count (Recall Count)8–12 itemsProvides enough potential matches for the reranking model while maintaining recall relevance.
Similarity threshold (Similarity Threshold)Calibrated by measurementBalances recall and precision based on actual query performance and data distribution.
Rerank result count (Rerank Return Count)3–5 itemsFocuses on the most relevant results, reduces the processing load on downstream large language models, and improves response speed.
PARSE_FILE_TIMEOUT_SECONDS600 secondsAddresses long parsing times for large PDF or Word documents, preventing parsing timeouts.

Common Pitfalls

  • Uploaded CSV files display garbled content, especially Chinese characters. This typically occurs when the CSV file's encoding format does not match the system's default encoding (e.g., a GBK-encoded file read as UTF-8).
  • Retrieval results include numerous irrelevant historical document versions. This happens when the knowledge base lacks effective version management or filtering, leading to the recall of outdated data.
  • After importing large documents, knowledge base disk space usage increases rapidly, exceeding expectations. This may result from excessively fine-grained document chunking, generating too many redundant text blocks and vectors, or from original files not being cleaned up promptly.

Verification Steps

  • Select several representative ADC quality documents, upload and chunk them. Check if the chunked text blocks are semantically complete and without obvious truncation, especially near tables and figure captions.
  • For queries targeting specific batch numbers or test items, observe whether recall results include the latest batch report for that batch number and relevant test data, and if historical versions are effectively filtered.
  • Simulate a user query for "purity test results for a certain batch product." Check if recall results accurately return text segments containing purity percentages and corresponding units (e.g., %).
  • Upload an Excel document with complex tables. Check if chunking results maintain the internal data relationships of the table, for example, if test items and corresponding values are in the same or adjacent text blocks.

Note: The values provided are common starting points. They should be measured against your own samples for optimal performance.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.