Document Parsing and Chunking for Antibody-Drug Conjugate (ADC) Regulatory Submissions

Antibody-Drug Conjugate (ADC) regulatory submission data involves multimodal information. This primarily includes preclinical study reports, clinical

Data Characteristics for ADC Regulatory Submissions

Antibody-Drug Conjugate (ADC) regulatory submission data involves multimodal information. This primarily includes preclinical study reports, clinical trial reports, and manufacturing process documents. Data sources are diverse, encompassing laboratory analysis reports, clinical data submitted by Contract Research Organizations (CROs), and internal manufacturing batch records. The update frequency typically aligns with drug development phases; for example, clinical trial data updates periodically, and manufacturing process data updates after process optimization. Document structures are complex, often in PDF format as CTD (Common Technical Document) modules. These contain numerous nested tables, figures (such as mass spectrometry and chromatography), and biological sequence information. Fields and units are highly specialized, for instance, pharmacokinetic (PK) parameters like Cmax and AUC, toxicology parameters like NOAEL, and antibody concentration units like μg/mL and drug-to-antibody ratio (DAR). Precise identification of these is critical.

Constraints from "Document Parsing and Chunking"

The complexity of ADC submission data imposes specific requirements on document parsing and chunking. First, nested tables and figures in PDFs mean traditional text-based chunking strategies may lose critical information. This necessitates enhanced table structure recognition and image OCR capabilities. Second, accurate identification of specialized fields and units is crucial for subsequent information extraction and knowledge graph construction. Parsing errors can lead to data correlation issues. Frequent updates mean the knowledge base must support incremental updates and version management. Chunking strategies need to consider how to efficiently identify and integrate new and old data. Additionally, the hierarchical structure of CTD modules requires chunking to preserve the document's logical hierarchy, for example, maintaining associations between multiple sub-reports of the same study to avoid semantic fragmentation.

Configuration Settings

Configuration ItemRecommended ValueRationale
Chunk size (Chunk Length)800–1200 charactersBalances semantic completeness and retrieval efficiency. Avoids excessive length leading to information redundancy and insufficient length causing context loss.
Chunk overlap (Chunk Overlap)100–200 charactersEnsures semantic continuity at chunk boundaries, especially in descriptive text for tables and figures.
File Type Whitelistpdf, docx, xlsx, txtCovers the main file formats for ADC submission documents, ensuring all relevant documents can be processed.
Table Parsing Mode (Table Parsing Mode)Structured ExtractionFor the numerous pharmacological and toxicological data tables in ADC reports, ensures data is saved in a structured format.
PARSE_FILE_TIMEOUT_SECONDS600 secondsAccounts for the long parsing time of large CTD files, providing sufficient time to prevent parsing interruptions.
OCR_ENABLEDTrueADC documents often contain scanned pages or figures. Enabling OCR identifies text information within images.

Common Pitfalls

  • Symptom: Table data is missing or misaligned in knowledge base retrieval results. Reason: The correct table parsing mode was not enabled or configured, causing table content to be treated as plain text and structural information to be lost.
  • Symptom: After uploading Feishu online spreadsheets, the knowledge base fails to correctly recognize their content. Reason: The system's default file parser may not support real-time synchronization or specific API interfaces for online documents. This requires using webhooks or exporting to a standard file format before uploading.
  • Symptom: Retrieved information after chunking has incomplete context, leading to illogical AI responses. Reason: The Chunk size (Chunk Length) setting is too small, or the document's chapter logic was not considered, causing semantically related content to be split into different chunks.

How to Verify Configuration

  • Upload a typical ADC submission PDF. Check if the parsed text chunks completely retain descriptive text and key data from tables and figures.
  • For documents containing specialized terminology and units, perform keyword searches. Verify if text chunks containing this information are accurately recalled.
  • Upload an updated document. Check if the knowledge base can identify incremental content and effectively associate it with existing knowledge.
  • Randomly select several chunks and manually review their content. Evaluate their semantic completeness, ensuring each chunk contains a relatively independent and meaningful unit of information.

The values provided are common starting points and should be measured against the reader's own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.