Document Parsing and Chunking for CAR-T Cell Therapy Registration Submissions

CAR-T cell therapy registration submission data involves multidisciplinary information. Core data sources include clinical trial reports (IND/NDA

Data Characteristics

CAR-T cell therapy registration submission data involves multidisciplinary information. Core data sources include clinical trial reports (IND/NDA stages), manufacturing process documents (CMC section), non-clinical study reports (pharmacology and toxicology), and quality control records. These data typically exist in formats such as PDF, Word, and Excel. Clinical trial data updates frequently, especially during ongoing multi-center clinical trials where data accumulation and analysis are continuous. Document structures are complex, containing numerous tables, figures, biological sequence information, and specialized terminology. Fields involve dose response, adverse event grading (CTCAE standards), cell expansion rates, viral vector titers, etc. Units include cell counts (cells/kg), gene copy numbers (GC/cell), and drug concentrations (μg/mL), reflecting specific biomedical measurement methods.

Constraints on Document Parsing and Chunking

The complexity of CAR-T cell therapy data imposes multiple constraints on document parsing and chunking. First, the presence of numerous figures and tables requires parsers to have high-accuracy recognition and structured extraction capabilities. Traditional text-stream-based splitting methods easily lose table context. Second, the high density of specialized terms and abbreviations (e.g., CAR, CR, PR) requires ensuring term integrity during chunking to avoid semantic deviation due to sentence breaks. Third, clinical reports have strict chapter logic, such as methodology, results, and discussion. Chunking must respect the original logical structure, avoiding the confusion of unrelated paragraphs. Fourth, high data update frequency requires parsing processes to support incremental updates and version management, ensuring knowledge base timeliness. Additionally, when dealing with biological sequences and gene editing content, specific pattern recognition and entity extraction accuracy are more demanding.

Configuration Settings

Configuration ItemSuggested ValueRationale
Chunk size800–1200 charactersBalances semantic completeness and retrieval efficiency. Avoids excessive length that introduces irrelevant information or excessive brevity that loses context.
Maximum Paragraph Depth5Accommodates the multi-level heading structure of clinical trial reports and CMC documents, preventing over-splitting of sub-chapter content.
Table Parsing ModeStructured ExtractionCAR-T data contains a large volume of tabular data, requiring precise extraction of row, column relationships, and cell content.
OCR Recognition AccuracyHighAddresses scanned historical documents or image-based figures, ensuring no loss of textual information.
Entity Recognition ListCAR-TCell type、target、adverse eventsPre-defines key biomedical entities, enhancing semantic accuracy of chunks and aiding subsequent retrieval.
Splitting StrategyBased on Title and Paragraph Length MixedPrioritizes maintaining content integrity under headings, supplemented by length limits for fine-grained splitting.

Three Common Pitfalls

  • Observation: Knowledge base content lacks important table data, or table content is incorrectly parsed as plain text. Reason: Table Parsing Mode (Table parsing mode) is not set to structured extraction, or the parser fails to correctly identify the boundaries and structure of complex tables.
  • Observation: During retrieval, some specialized terms are truncated, leading to incomplete semantics or inaccurate recall. Reason: Chunk size (Chunk length) is set too small, causing sentences containing long compound terms to be forcibly split.
  • Observation: After uploading large PDF documents, the parsing task remains unresponsive for a long time or reports a timeout error. Reason: PARSE_FILE_TIMEOUT_SECONDS parameter is set too low, or the document contains too many embedded images, leading to OCR taking longer than expected.

How to Verify Configuration

  • Randomly select 5-10 typical documents and inspect their parsed knowledge chunks to confirm whether table data is complete and structured.
  • Examine knowledge chunks containing specific specialized terms (e.g., CD19 CAR-T, cytokine release syndrome) in the knowledge base to assess their semantic completeness.
  • Upload a comprehensive clinical study report containing multi-level headings, figures, and long paragraphs. Observe the completion time and results of the parsing task to ensure no timeout errors.
  • Compare the original document with the parsed knowledge chunks. Check if the boundaries of key sections (e.g., Methods, Results, Discussion) align with expectations, with no unreasonable merging or splitting.

The values provided are common starting points and should be measured against the reader's own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.