Document Parsing and Chunking for Antibody-Drug Conjugate (ADC) Products

Antibody-Drug Conjugate (ADC) product data originates from various sources. These primarily include official drug inserts, clinical trial reports

Understanding ADC Product Data

Antibody-Drug Conjugate (ADC) product data originates from various sources. These primarily include official drug inserts, clinical trial reports, research papers, patent literature, and public databases from drug regulatory agencies. Document update frequency depends on drug development, approval, and post-market surveillance progress. For example, clinical trial data updates regularly as trial phases advance, and drug inserts may revise based on adverse event monitoring or expanded indications.

ADC product data typically features highly structured information. This includes drug components, mechanisms of action, pharmacokinetics, pharmacodynamics, indications, dosage and administration, adverse reactions, and contraindications. Specifically, component information covers the precise molecular structures and CAS numbers of antibodies, linkers, and cytotoxic drugs. Pharmacokinetic data includes half-life and clearance rates. Pharmacodynamic data involves quantitative metrics such as binding affinity and tumor inhibition rates.

Units commonly used are milligrams (mg) or milligrams per kilogram (mg/kg) for dosage, micrograms per milliliter (µg/mL) or nanomoles (nM) for concentration, hours (h) or days (d) for time, and milliliters (mL) for volume.

Constraints on Document Parsing and Chunking

The diverse and structured nature of ADC product documents imposes specific requirements on parsing and chunking.

First, the prevalence of tables and figures, especially in clinical trial results and pharmacokinetic data, demands robust table content recognition and structured extraction capabilities to prevent data misalignment or loss.

Second, frequent molecular structures and chemical names necessitate support for specialized chemical entity recognition. This prevents critical information from being incorrectly tokenized or truncated.

Third, varying update frequencies across different document sources—such as rapid updates in clinical research progress—require the knowledge base to efficiently identify and process incremental data, ensuring information timeliness.

Fourth, ADC product descriptions often contain complex technical terms and abbreviations. Chunking must preserve contextual integrity to avoid semantic loss due to over-segmentation.

Finally, numerical values and units in pharmacokinetic and pharmacodynamic data are tightly coupled. Chunking must maintain the proximity of values to their corresponding units; otherwise, it could lead to data misinterpretation or calculation errors.

Configuration Recommendations

Configuration ItemRecommended ValueRationale
Chunk size500–800 charactersBalances the contextual integrity of complex technical terms with retrieval efficiency, avoiding information redundancy from overly long texts.
Chunk Overlap Length100–150 charactersEnsures continuity of information across segments, particularly when describing mechanisms of action and pharmacokinetic processes.
File Type Whitelistpdf, docx, xlsx, txt, jsonCovers common ADC product document formats, especially PDF and DOCX for inserts and reports, and XLSX for experimental data.
Table Recognition ModeAutomatic Recognition And StructuredAddresses the large number of tables in clinical trial reports and pharmacokinetic data, ensuring accurate data extraction.
OCR语言Chinese, EnglishHandles global research literature and drug inserts, ensuring accurate recognition of multilingual content.
ParsingTimeout600 secondsAccounts for potentially long parsing times for large clinical reports or documents containing many images/tables.

Common Pitfalls

  • Parsing logs show "slow operation xxxxms" due to slow MongoDB response. This indicates large file sizes or complex tables causing the parser to exceed expected processing time.
  • Uploading a DOCX file results in "Invalid image file" for image content. This means the parser failed to correctly identify and process embedded image data, or image information was not extracted or converted to text.
  • Markdown table content in model output is truncated. This happens when the chunk length is set too small, causing a single table to be split across multiple segments, or a single text segment exceeds the model's input length limit.

Verification Steps

  • Upload a PDF drug insert containing complex tables. Check if the parsed text fully retains table structures and data, and perform random sampling comparisons.
  • Upload a patent document with molecular structure diagrams. Verify that image descriptions or related text are accurately extracted, and confirm that critical chemical entities are not incorrectly segmented.
  • Upload an updated version of a clinical trial report. Check if incremental data is correctly identified and updated in the knowledge base, and verify that key pharmacokinetic parameters maintain their numerical and unit associations.

Note: The values provided are common starting points and should be measured against your own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.