Document Parsing and Chunking for Solid Tumor Clinical Trial Pre-screening

Solid tumor clinical trial pre-screening data originates from sponsor-provided documents. These include Protocols, Investigator's Brochures (IB)

Data Characteristics

Solid tumor clinical trial pre-screening data originates from sponsor-provided documents. These include Protocols, Investigator's Brochures (IB), Informed Consent Forms (ICF), patient medical records, medical imaging reports, and genetic testing reports. Most documents are in PDF format; some may be scanned images. Protocols and IBs have a low update frequency, typically updated during trial initiation or major protocol amendments. Patient medical records and test reports are highly time-sensitive, generated in real-time with patient visits and test results. Document structures are complex, containing numerous tables, figures, and nested lists. Text content involves medical terminology, units of measurement (e.g., mg/kg, mm, %), specific disease staging (e.g., TNM staging), and efficacy evaluation criteria (e.g., RECIST 1.1).

Constraints on Document Parsing and Chunking

The complex structure of solid tumor clinical trial documents demands high-quality document parsing. This is especially true for correctly identifying and extracting content from tables and nested lists. Scanned PDFs require high-quality OCR for complete text content. Highly time-sensitive patient data requires rapid parsing and processing to avoid information delays. Extensive medical terminology and professional abbreviations mean simple text segmentation may fail to capture key concepts accurately. Structured information, such as disease staging and efficacy standards, must maintain semantic integrity during chunking. This prevents splitting across different knowledge chunks, which would affect subsequent accurate retrieval. Identifying units of measurement is crucial for evaluating numerical conditions; the parser must distinguish between values and units.

Configuration Settings

Configuration ItemRecommended ValueRationale
Chunk size (Chunk Length)600–800 characters (characters)Balances context completeness and retrieval efficiency. Avoids excessively long text diluting key information and excessively short text losing context.
Chunk Overlap Length (Chunk Overlap Length)100–150 characters (characters)Ensures semantic coherence at chunk boundaries. This is especially important for critical medical concepts or tables spanning pages.
PARSE_FILE_TIMEOUT_SECONDS600 seconds (seconds)Handles large PDFs or those with complex tables. Requires a longer parsing time to prevent timeout errors.
chunk_strategyBy title and paragraphPrioritizes preserving the document's logical structure. Ensures related information stays within the same chunk, especially suitable for medical protocols.
table_parsing_modeAdvanced modeAddresses the large number of complex tables in solid tumor clinical trial documents. Ensures accurate extraction and structuring of table content.
OCR_ENABLEDTrueEnsures scanned medical records and reports are recognized and parsed. This obtains complete text content.

Common Pitfalls

  • Content is not recognized after uploading a PDF with a digital signature. The digital signature may interfere with the document parsing library's ability to correctly read the file structure.
  • Table content is lost or garbled when parsing complex tables. Default table parsing modes may not handle nested tables, merged cells, or irregular table lines.
  • Parsing large PDF files results in timeout errors. The default file parsing timeout is insufficient to process all file content.

Verification Steps

  • Select a solid tumor clinical trial protocol PDF containing complex tables and multiple pages. Upload it to the knowledge base. Check if the parsed knowledge chunks fully include table data.
  • Upload a scanned patient medical record PDF. Observe if its text content is correctly recognized, with no obvious garbling or omissions.
  • Check knowledge chunks related to disease staging (e.g., TNM staging) or efficacy criteria (e.g., RECIST 1.1) in the knowledge base. Confirm that the descriptions of these concepts maintain semantic integrity and are not improperly split.

The values provided are common starting points. Measure them against your own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.