Data Characteristics
CAR-T cell therapy regulatory documents originate from the National Medical Products Administration (NMPA) regulations and technical guidelines, as well as internal Standard Operating Procedures (SOPs) and Quality Management System documents from medical institutions and pharmaceutical companies. These documents are primarily in PDF format. They contain extensive specialized terminology, charts, flowcharts, and complex approval requirements. Updates typically occur quarterly or annually, driven by policy adjustments and clinical practice advancements. Urgent updates may occur for significant safety or efficacy changes. Documents feature a rigorous, hierarchical structure with clear chapter titles and clause numbers. Fields and units include cell preparation batch numbers, patient identifiers, formulation dosages, culture times, and quality control indicators (e.g., cell viability, purity, gene copy number). Units are precise, such as micrograms, milliliters, and hours.
Constraints on "Document Parsing and Chunking"
The specialized nature, structured characteristics, and update frequency of CAR-T cell therapy regulatory documents impose specific requirements on document parsing and chunking. Complex charts and flowcharts mean that plain text extraction may lose critical information, necessitating OCR or image parsing capabilities. Strict chapter numbering and clause divisions require chunking to preserve the original logical structure, preventing semantic fragmentation. Tabular data, such as QC indicators, needs special handling to maintain contextual integrity. Regular policy updates demand knowledge base support for incremental updates and version management to ensure retrieved information is current and valid. Identifying specific fields and units helps the subsequent Q&A system accurately understand user queries, avoiding errors due to unit confusion.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Chunk Length) | 500–800 characters (characters) | Balances semantic completeness with retrieval efficiency. Avoids excessively long chunks that introduce redundancy or overly short chunks that lose context. |
Chunk Overlap Length (Chunk Overlap Length) | 50–100 characters (characters) | Ensures contextual continuity at chunk boundaries, reducing loss of key information due to sentence breaks. |
OCR_ENABLED | True | CAR-T documents often contain charts and flowcharts. Enabling OCR helps extract non-textual information. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds (seconds) | Processing large PDF documents or complex OCR tasks can be time-consuming. Increases the timeout limit. |
CHUNK_STRATEGY | By Title and Paragraph | Prioritizes chunking based on document chapter titles and paragraph structure to maintain logical integrity. |
TABLE_EXTRACTION_MODE | Structured Text | Preserves table row and column relationships for accurate understanding of tabular data during retrieval. |
Common Pitfalls
PARSE_FILE_TIMEOUTerrors occur when parsing large PDF files because the default timeout is insufficient for documents with many images or complex layouts.- After uploading XLSX files containing critical tabular data, table content is missing or garbled in Q&A results because the table parsing mode is not set to preserve structured information.
- When users ask about a specific version of a regulation, the system returns inaccurate answers because the knowledge base lacks version management or the incremental update strategy fails to handle old document versions effectively.
Verification Steps
- Upload a PDF document containing a complex flowchart. Verify that text within the flowchart is correctly extracted and included in the parsed chunks.
- Upload an XLSX document containing multiple rows of quality control tables. Verify that table data is presented in a structured format (e.g., Markdown table) within the parsed chunks.
- Upload a new version of an SOP document. Then, retrieve clauses related to that SOP. Confirm that Q&A results are from the latest version and can distinguish between old and new versions.
- Randomly select specialized terms or key sentences from the document for retrieval. Verify that the recalled chunks are semantically complete and traceable to the corresponding location in the original document.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.