Document Parsing and Chunking for Solid Tumor Regulatory Submission Preparation

Solid tumor regulatory submission documents typically include clinical trial protocols, investigator brochures, clinical study reports (CSRs)

Data Characteristics

Solid tumor regulatory submission documents typically include clinical trial protocols, investigator brochures, clinical study reports (CSRs), statistical analysis plans (SAPs), medical writing documents, and safety reports. These documents are primarily in PDF, Word, or scanned image formats. They have complex structures, often containing numerous charts, medical images, biomarker data, and specialized terminology. Data update frequency is low, mainly concentrated during the release of interim reports at different clinical trial stages and the final submission phase. Documents contain many internal fields such as patient ID, tumor stage, treatment regimen, efficacy evaluation criteria (RECIST standards), adverse event (AE) grades, and laboratory test results. These often come with specific medical units (e.g., mg/kg, mm, U/L).

Constraints Imposed by These Characteristics on Document Parsing and Chunking

The complex and multi-source nature of solid tumor submission documents demands high-quality document parsing. Large numbers of charts and scanned images require high-precision OCR capabilities to ensure complete text extraction. Dense specialized terminology and abbreviations can lead to poor performance from general-purpose tokenizers, necessitating customized dictionary support. Medical images and biomarker data usually appear as pictures or tables; their semantic parsing requires contextual understanding, as simple text chunking cannot capture their deeper meaning. The strict correspondence between fields and units requires chunks to maintain data associations for accurate retrieval and question-answering. Low update frequency means the knowledge base remains stable once parsed, making initial parsing accuracy critical.

Configuration Settings

Configuration ItemRecommended ValueRationale
Chunk size800–1200 charactersSolid tumor documents have high content density. Moderately increasing chunk length helps retain context and reduces information fragmentation.
Chunk Overlap Length100–200 charactersEnsures semantic continuity between adjacent chunks, especially for cross-paragraph specialized terms or data descriptions.
OCR_ENGINE_TYPEPaddleOCR or TesseractHigh-precision OCR engines are needed for scanned images and complex charts. Benchmark different engines for medical image recognition performance.
CHUNK_STRATEGYChunk by Title combined with by fixed lengthPrioritizes using the document's chapter structure, then supplements with fixed-length splitting, balancing structure and content completeness.
MAX_FILE_SIZE_MB100 MBRegulatory submission documents often contain many images and charts, leading to large PDF file sizes. File size limits need appropriate relaxation.
PARSE_TIMEOUT_SECONDS600 secondsParsing large PDF documents can be time-consuming. Increase the timeout to prevent parsing interruptions.

Common Mistakes

  • Garbled or missing content in parsing results: This usually occurs when the OCR engine's recognition capability is insufficient for specific fonts or poor-quality scanned medical images, or when the OCR language pack is not configured correctly.
  • Model cannot accurately answer questions involving charts or table data: This happens when document parsing fails to effectively extract text information from charts or does not structure table data, leading to a lack of this content in the knowledge base.
  • System becomes unresponsive or returns a 504 Gateway Timeout error after uploading large PDF files: This may relate to a low PARSE_TIMEOUT_SECONDS configuration, where large file parsing time exceeds the system's default or configured timeout limit.

How to Verify Configuration

  • Randomly select multiple solid tumor regulatory submission documents of different types. Check if the parsed text content is complete and free of garbled characters, comparing it against the original documents.
  • For documents containing charts and tables, verify if chart titles, legends, and key data within tables are correctly extracted in the parsing results.
  • Conduct question-answering tests on the knowledge base. Questions should cover different sections and data types (text, tabular data) within the documents. Evaluate the accuracy and completeness of the model's answers, ensuring key fields and units are correctly identified.
  • Check system logs to ensure no significant parsing failures or timeout errors occur, especially for large submission documents.

The values provided are common starting points. Measure their effectiveness against your own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.