Data Characteristics
CAR-T cell therapy product data primarily originates from clinical trial reports, drug labels, regulatory approval documents, academic papers, and patent literature. Document update frequency aligns with product lifecycles and regulatory requirements. For example, clinical trial results might update monthly or quarterly, while drug labels undergo annual revisions after market approval or supplemental updates when safety information changes. Document structures are typically highly standardized, following templates like ICH E3 for clinical study reports or FDA drug labels, featuring fixed sections and data presentation formats. Fields involve complex biological and medical terminology, such as target expression, cell expansion rates, pharmacokinetic parameters, and adverse event incidence. Units include cell concentration (e.g., cells/kg), dosage (e.g., mg/kg), time (e.g., days, months), and various biomarker concentrations (e.g., pg/mL).
Constraints Imposed by These Characteristics on Document Parsing and Chunking
The highly standardized structure of CAR-T cell therapy product documents makes a hybrid chunking strategy, based on section titles and semantic content, more effective. Traditional text chunking can lead to truncation of critical information or loss of context due to the large volume of specialized terminology and complex data tables. This necessitates more refined paragraph boundary identification and table content extraction. The periodic nature of document updates requires the knowledge base to support incremental updates and version management, ensuring the timeliness of parsed content. Additionally, synonyms, abbreviations, or inconsistent units across different source documents pose challenges for entity recognition and standardization after parsing. For instance, the CD19 target might appear in various forms across different literature, requiring unified handling. Clinical trial data often includes numerical ranges and unit conversions, demanding that the parser handles complex numerical information.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk Length | 800–1200 characters | Balances the contextual relevance of specialized terminology in CAR-T documents with the integrity of large paragraphs, preventing critical information from being split. |
Overlap Length | 100–150 characters | Ensures sufficient contextual overlap between adjacent chunks, especially when describing complex biological mechanisms or clinical outcomes. |
Parsing Mode | Prioritize structured parsing, supplemented by semantic chunking | Leverages the standardized chapter structure of CAR-T documents, independently processing tables and figure captions. |
Table Parsing Strategy | Identify table boundaries, extract data by row or column, and retain header information | Ensures accuracy of numerical and descriptive information in tables, such as clinical data and adverse event statistics. |
Timeout (PARSE_FILE_TIMEOUT_SECONDS) | 600 seconds | Accounts for potentially long parsing times for large clinical trial reports (e.g., PDFs over 100MB). |
Image OCR Recognition | Enabled | Many CAR-T documents contain critical flow cytometry charts, histopathological images, or structural diagrams, requiring text extraction from images. |
Three Common Pitfalls
- The parsing status remains "parsing" for an extended period or ultimately displays "parsing failed." This can occur if the uploaded PDF file is too large or its content is overly complex, causing the parser to exceed the
PARSE_FILE_TIMEOUT_SECONDSsetting. - Critical clinical trial data table content is missing or disorganized in the parsing results. This usually happens when the table parsing strategy is not optimized for the complex table structures found in CAR-T documents, failing to correctly identify headers or data cells.
- AI responses show misunderstandings of specialized terminology or lack context. This often results from an improper
Chunk Lengthsetting, leading to related biological concepts or treatment plan descriptions being inappropriately split across different chunks.
How to Verify Configuration
- Upload and parse multiple CAR-T documents from different sources (e.g., FDA labels, clinical trial reports) and check if their parsing status uniformly displays "ready."
- Randomly select parsed documents and verify that the chunks for key sections (e.g., "Mechanism of Action," "Clinical Studies," "Adverse Reactions") are complete and semantically coherent, especially confirming that table data is correctly extracted.
- Perform tests on the parsed knowledge base involving specialized terminology and numerical queries to assess the accuracy and completeness of recall results. Examples include querying "complete response rate for CD19 positive patients" or "recommended dosage of tisagenlecleucel."
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.