Document Parsing and Chunking for Small Molecule Drug Registration Dossiers

Small molecule drug registration dossiers typically originate from experimental reports, test data, clinical study reports, quality standards, and

Data Characteristics

Small molecule drug registration dossiers typically originate from experimental reports, test data, clinical study reports, quality standards, and manufacturing process documents generated during drug research, development, and production. These documents have a low update frequency, primarily during new drug applications, supplemental applications, or periodic reports. Document structures are highly standardized, adhering to regulatory requirements such as ICH and NMPA's Common Technical Document (CTD) format. Common document types include PDF-formatted lab reports, chromatograms, and research reports, as well as Word or Excel-formatted application forms. Data fields include batch numbers, test items, test results, units (e.g., mg/mL, °C, pH value), test methods, and acceptance criteria. Chromatogram data (e.g., HPLC, NMR) is embedded as images.

Constraints Imposed by These Characteristics on Document Parsing and Chunking

The standardized nature of small molecule drug dossiers requires precise identification and extraction of structured information during document parsing, with particular sensitivity to key data in tables and chromatograms. Embedded images (e.g., chemical structures, chromatograms) in PDF documents are core information carriers. Document parsing must handle both text and image content, extracting searchable metadata or textual descriptions from them. The low document update frequency means initial parsing requires significant resources to ensure accuracy, with less pressure for subsequent incremental updates. The standardization of fields and units requires chunking to effectively associate values with units to avoid misinterpretation. Chunking strategies must consider the logical integrity of CTD modules, ensuring each chunk independently provides meaningful context. For example, a chunk should contain a complete experimental method description or all results for a specific test item.

Configuration Guidelines

Configuration ItemRecommended ValueRationale
Chunk size800–1200 charactersBalances the integrity of experimental report paragraphs for small molecule drugs with retrieval granularity, preventing truncation of critical descriptions or data tables.
Chunk Overlap Length100–200 charactersEnsures contextual continuity, especially in tables and chromatogram descriptions spanning pages or paragraphs.
Yes noEnabledImage tabletsOCRYesSmall molecule drug dossiers contain numerous embedded chromatograms and chemical structures. OCR is essential for extracting text from this information.
OCR语言Chinese, EnglishDossiers may contain both Chinese descriptions and English professional terminology or compound names.
Table Parsing ModeStructuredAccurately identifies and extracts values and units from tables such as lab reports and stability data, preserving their structural relationships.
Recall countTop 5 entriesRegistration dossier queries typically require fewer, highly relevant, and precise results, avoiding interference from irrelevant information.

Common Pitfalls

  • Image content is missing or incomplete in parsing results. The symptom is that recall results only contain text descriptions and lack image information. This occurs because Yes noEnabledImage tabletsOCR is not enabled or OCR recognition rates are insufficient, preventing embedded chromatograms or chemical structures from being converted into searchable text.
  • When querying specific data, the returned results contain a large amount of irrelevant context. The symptom is that recalled chunks are too large or contain multiple unrelated topics. This occurs because Chunk size is set too long, failing to effectively distinguish between different experiments or test items.
  • Table data query results are inaccurate. The symptom is that values are misaligned with corresponding test items or units. This occurs because Table Parsing Mode is not set to structured mode, causing table content to be flattened and losing row-column relationship information.

How to Verify Correct Configuration

  • Select a typical small molecule drug dossier PDF document containing complex tables and embedded chromatograms. After uploading and parsing, randomly select multiple chunks from the knowledge base. Check if their content includes text recognized by OCR from chromatograms and structured data from tables.
  • Perform a search for specific experiment project names or compound structure names. Check if the recalled chunks completely include all descriptions, data, and relevant chromatogram information for that project, and evaluate the coherence of the context.
  • Simulate a query for a specific test result for a batch (e.g., "Batch XXX content"). Verify if the numerical value in the recalled result accurately corresponds to the correct test item and unit, and confirm the page number information of its source document.

Note: The values provided are common starting points. They should be measured against your own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.