Data Characteristics
Solid tumor regulatory submission documents typically include clinical trial protocols, investigator brochures, clinical study reports (CSRs), statistical analysis plans (SAPs), medical writing documents, and safety reports. These documents are primarily in PDF, Word, or scanned image formats. They have complex structures, often containing numerous charts, medical images, biomarker data, and specialized terminology. Data update frequency is low, mainly concentrated during the release of interim reports at different clinical trial stages and the final submission phase. Documents contain many internal fields such as patient ID, tumor stage, treatment regimen, efficacy evaluation criteria (RECIST standards), adverse event (AE) grades, and laboratory test results. These often come with specific medical units (e.g., mg/kg, mm, U/L).
Constraints Imposed by These Characteristics on Document Parsing and Chunking
The complex and multi-source nature of solid tumor submission documents demands high-quality document parsing. Large numbers of charts and scanned images require high-precision OCR capabilities to ensure complete text extraction. Dense specialized terminology and abbreviations can lead to poor performance from general-purpose tokenizers, necessitating customized dictionary support. Medical images and biomarker data usually appear as pictures or tables; their semantic parsing requires contextual understanding, as simple text chunking cannot capture their deeper meaning. The strict correspondence between fields and units requires chunks to maintain data associations for accurate retrieval and question-answering. Low update frequency means the knowledge base remains stable once parsed, making initial parsing accuracy critical.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size | 800–1200 characters | Solid tumor documents have high content density. Moderately increasing chunk length helps retain context and reduces information fragmentation. |
Chunk Overlap Length | 100–200 characters | Ensures semantic continuity between adjacent chunks, especially for cross-paragraph specialized terms or data descriptions. |
OCR_ENGINE_TYPE | PaddleOCR or Tesseract | High-precision OCR engines are needed for scanned images and complex charts. Benchmark different engines for medical image recognition performance. |
CHUNK_STRATEGY | Chunk by Title combined with by fixed length | Prioritizes using the document's chapter structure, then supplements with fixed-length splitting, balancing structure and content completeness. |
MAX_FILE_SIZE_MB | 100 MB | Regulatory submission documents often contain many images and charts, leading to large PDF file sizes. File size limits need appropriate relaxation. |
PARSE_TIMEOUT_SECONDS | 600 seconds | Parsing large PDF documents can be time-consuming. Increase the timeout to prevent parsing interruptions. |
Common Mistakes
- Garbled or missing content in parsing results: This usually occurs when the OCR engine's recognition capability is insufficient for specific fonts or poor-quality scanned medical images, or when the OCR language pack is not configured correctly.
- Model cannot accurately answer questions involving charts or table data: This happens when document parsing fails to effectively extract text information from charts or does not structure table data, leading to a lack of this content in the knowledge base.
- System becomes unresponsive or returns a
504 Gateway Timeouterror after uploading large PDF files: This may relate to a lowPARSE_TIMEOUT_SECONDSconfiguration, where large file parsing time exceeds the system's default or configured timeout limit.
How to Verify Configuration
- Randomly select multiple solid tumor regulatory submission documents of different types. Check if the parsed text content is complete and free of garbled characters, comparing it against the original documents.
- For documents containing charts and tables, verify if chart titles, legends, and key data within tables are correctly extracted in the parsing results.
- Conduct question-answering tests on the knowledge base. Questions should cover different sections and data types (text, tabular data) within the documents. Evaluate the accuracy and completeness of the model's answers, ensuring key fields and units are correctly identified.
- Check system logs to ensure no significant parsing failures or timeout errors occur, especially for large submission documents.
The values provided are common starting points. Measure their effectiveness against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.