Document Parsing and Chunking for CAR-T Cell Therapy Clinical Trial Pre-screening

CAR-T cell therapy clinical trial data primarily originates from sponsor documents submitted by clinical research organizations and public information

Data Characteristics

CAR-T cell therapy clinical trial data primarily originates from sponsor documents submitted by clinical research organizations and public information released by regulatory bodies. These documents have a relatively low update frequency, typically released with trial progress or reporting cycles, such as annual reports, interim analysis reports, and final study reports. Document structures are complex, including protocols, informed consent forms, case report forms (CRFs), investigator brochures (IBs), ethics committee approval letters, and various laboratory test reports. Field types are diverse, covering patient basic information, diagnostic criteria, treatment plans, dose adjustments, adverse events, efficacy evaluation indicators (e.g., complete response rate CR, partial response rate PR), cell preparation details, gene editing information, and biomarker data. Units involve dosage (e.g., cells/kg), time (e.g., months, weeks), percentages (e.g., tumor burden reduction %), and various laboratory test values (e.g., pg/mL, IU/mL).

Constraints Imposed by These Characteristics on Document Parsing and Chunking

The complexity of CAR-T clinical trial documents demands high-quality document parsing. Multiple sources lead to inconsistent document formats, potentially mixing PDFs, Word documents, structured tables, and even scanned images. This requires effective identification and information extraction across various file types. The specialized and diverse nature of fields, especially data with specific biological or medical abbreviations, requires the parser to have a high degree of semantic understanding. This prevents misidentification of key indicators as ordinary text. For example, CR could refer to complete response or other medical abbreviations, requiring context for accurate judgment. Low update frequency means initial parsing accuracy is critical; errors can have long-term impacts on subsequent pre-screening results. Furthermore, the presence of numerous tables and nested lists makes it difficult for traditional rule-based or simple text segmentation methods to accurately capture data relationships. This can lead to fragmentation of key information, affecting the accuracy of subsequent vectorization and retrieval.

Configuration Settings

Configuration ItemRecommended ValueRationale
UPLOAD_FILE_MAX_SIZE500 MBClinical trial documents often contain many charts and images, resulting in large file sizes. Sufficient upload capacity is necessary.
PARSE_FILE_TIMEOUT_SECONDS600 secondsProcessing large PDFs or Word documents with complex tables can be time-consuming. This prevents parsing timeouts.
Chunk size800–1200 charactersBalances contextual completeness with retrieval efficiency. Each chunk contains enough information for semantic understanding while reducing irrelevant information interference.
Chunk Overlap Length100–200 charactersEnsures information continuity at chunk boundaries, reducing key information fragmentation caused by chunking.
Max Chunks1000Limits the number of chunks generated per document. This prevents excessively long documents from consuming too many resources and ensures retrievable results are manageable.
Table Parsing ModeSmart ParsingTables in clinical documents are complex and nested. Intelligent parsing mode better identifies table boundaries and cell content, preserving data structure.

Three Common Mistakes

  • Key fields (e.g., adverse event grades, efficacy evaluation results) are empty after document parsing. This can happen if the parser fails to recognize specific medical terms or abbreviations, or if the table structure is too complex for data extraction.
  • A 504 Gateway Timeout error occurs when uploading large PDF files. This usually means PARSE_FILE_TIMEOUT_SECONDS is set too short, and the file parsing time exceeds the allowed limit.
  • Retrieval results contain many irrelevant chunks. This can happen if Chunk size is too long, causing individual chunks to contain excessive noise, diluting key semantics.

How to Confirm Proper Configuration

  • Select several typical CAR-T clinical trial documents, upload them, and check the parsed text content. Pay close attention to the completeness and accuracy of specialized terms, dosage units, and table data.
  • Perform keyword and semantic searches on the parsed documents. Observe whether the returned results include the chunks containing the target information, and evaluate the contextual completeness of the chunks.
  • Check system logs to confirm that no timeouts or other parsing errors occurred when processing various documents, especially for large or complex formats.
  • Verify key field extraction. For example, randomly select several adverse event reports from the parsing results. Confirm that their grades and descriptions match the original documents. Adjust Table Parsing Mode and Chunk size parameters accordingly.

The values given are common starting points and should be measured against your own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.