Data Characteristics
CAR-T cell therapy quality documents originate from drug registration applications, manufacturing batch records, quality standards, validation reports, and Standard Operating Procedures (SOPs). Regulatory requirements, process changes, and clinical trial progress influence the update frequency of these documents, which typically occurs in concentrated phases. Structurally, registration application documents are often large, multi-level PDF files with cross-references, containing numerous charts, flowcharts, and specialized terminology. Manufacturing batch records are mostly structured or semi-structured tabular data, recording parameters and results for key steps like cell expansion, transfection, and quality control. Quality standard documents strictly define product testing indicators, methods, and limits. Common fields include biological and molecular biological indicators such as cell count (units: cells/mL), viability (units: %), gene copy number (units: copies/μL), and viral titer (units: TU/mL).
Constraints Imposed by These Characteristics on "Document Parsing and Chunking"
The complexity of CAR-T cell therapy quality documents imposes specific requirements on document parsing and chunking. First, charts and flowcharts in multi-level PDF documents require the parser to accurately identify images and extract relevant text descriptions, preventing information loss. Second, structured tabular data in batch records requires the parser to correctly identify rows and columns and convert them into queryable structured text, ensuring accuracy in subsequent question-answering. The extensive presence of specialized terminology and abbreviations, such as "CAR," "lentivirus," and "GMP," necessitates maintaining contextual integrity during chunking to avoid semantic fragmentation due to over-segmentation. Furthermore, the phased nature of data updates means that incremental parsing and index reconstruction must be performed efficiently during batch updates to reflect the latest quality status. For critical quality indicators and test results, chunking must ensure that values and units are not separated, maintaining data integrity.
Configuration Recommendations
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Chunk Length) | 800–1200 characters | Balances specialized terminology and contextual integrity, preventing semantic loss from over-segmentation. |
Chunk Overlap Length (Chunk Overlap Length) | 100 characters | Ensures sufficient overlap between chunks to connect context and handle cross-chunk specialized terms. |
Parsing Mode | Smart Parsing | Addresses complex layouts in PDFs, including charts, flowcharts, and tables, improving information extraction accuracy. |
Table Content Handling | Structured Extraction | For tabular data in manufacturing batch records, ensures row and column information is not lost, facilitating querying. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Large registration application documents take longer to parse; prevents parsing failures due to timeouts. |
Character Cleaning Rules | Retain Special Symbols | Many biomedical terms contain special symbols and subscripts; retaining them aids accurate identification. |
Three Common Mistakes
- Symptom: After uploading a PDF, text or tabular data within some charts cannot be retrieved. Reason: The parsing mode is incorrectly selected, failing to effectively recognize text within images or table structures.
- Symptom: After uploading batch records, query results for critical parameters (e.g., cell viability) for a specific batch are incomplete, with values and units separated. Reason: The chunk length is set too small, causing values and units to be split during segmentation, losing integrity.
- Symptom: After file upload, the parsing node is unresponsive for an extended period or returns a
504 Gateway Timeouterror. Reason: ThePARSE_FILE_TIMEOUT_SECONDSparameter is set too low, and parsing large or complex documents exceeds the time limit.
How to Confirm Correct Configuration
- Upload a typical CAR-T quality document containing charts, tables, and specialized terminology. Check the completeness and accuracy of these elements in the parsed text.
- For a structured batch record document, query key numerical fields (e.g., cell count, viral titer) to verify that both the value and its unit are recalled together.
- Attempt to upload a complex PDF document with a file size close to the system limit. Observe whether the parsing process completes smoothly without timeout errors.
Note: The values provided are common starting points. They should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.