Document Parsing and Chunking for CAR-T Cell Therapy Pharmacovigilance

CAR-T cell therapy pharmacovigilance data primarily originates from clinical trial reports, real-world evidence (RWE) data, case reports, post-market

Data Characteristics

CAR-T cell therapy pharmacovigilance data primarily originates from clinical trial reports, real-world evidence (RWE) data, case reports, post-market surveillance reports, and regulatory guidelines. These documents are typically in PDF format, with some in Word or exported from structured databases. Data updates frequently, especially during new product launches and regulatory policy changes. Document structures are complex, containing extensive medical terminology, abbreviations, charts, and tables. Information includes patient demographics, CAR-T product details, adverse event (AE) descriptions, serious adverse event (SAE) reports, laboratory test results, concomitant medications, and treatment outcomes. Fields may include standardized medical codes (e.g., MedDRA codes) and free-text descriptions. Units involve dosage (e.g., cells/kg), time (e.g., days, weeks), and laboratory indicators (e.g., pg/mL, %). Field names and unit representations can vary across different source documents.

Constraints on Document Parsing and Chunking

The complex structure and high information density of CAR-T cell therapy documents demand advanced parsing capabilities. Large volumes of charts and tables require accurate extraction of key data, preventing information loss. The prevalence of medical terminology and abbreviations necessitates preserving contextual integrity during chunking to avoid semantic ambiguity caused by truncation. For example, a complete adverse event description might span multiple paragraphs or include tabular data. Mechanically chunking by paragraph or fixed character count can fragment critical information. High data update frequency requires the parsing and chunking process to support efficient incremental updates and version management, ensuring knowledge base timeliness. Diverse fields and units require parsed content to retain its original context during chunking, facilitating subsequent semantic understanding and retrieval. For clinical trial reports often hundreds of pages long, parser stability and resource consumption are key considerations.

Configuration Recommendations

Configuration ItemRecommended ValueRationale
UPLOAD_FILE_MAX_SIZE500 MBClinical trial reports can be large; ensure full upload capability.
PARSE_FILE_TIMEOUT_SECONDS600 secondsLarge file parsing is time-consuming; prevent timeout failures.
Chunk size800–1200 charactersBalances contextual completeness and retrieval efficiency; avoids overly long or short chunks.
Chunk Overlap Length100–200 charactersEnsures semantic continuity at chunk boundaries, especially for medical terms and descriptions.
maxContext3For complex medical texts like CAR-T, retaining more context aids understanding.
Recall countTop 5-8 entriesIncreases recall rate of relevant information, covering different angles of adverse event descriptions.

Common Pitfalls

  • Parsing hundreds of PDF pages fails, while parsing tens of pages succeeds. This typically occurs when PARSE_FILE_TIMEOUT_SECONDS is set too low, causing the parser to time out when processing large, complex documents.
  • Key laboratory indicator values or dosage information are missing from chunked recall results. This usually happens when Chunk size is set too small, leading to the fragmentation of tables or chart captions containing critical values, preventing the formation of complete semantic blocks.
  • Locally deployed pdf-marker throws a Cannot read properties of undefined error. This issue may relate to compatibility between the pdf-marker version and the FastGPT version, or incorrect environment dependency installation.

Verification Steps

  • Upload a CAR-T clinical trial report containing complex tables and charts. Check if the parsed text content fully retains table data and chart descriptions.
  • For a typical adverse event report document, adjust Chunk size and Chunk Overlap Length. Then, use the knowledge base preview function to verify if the chunked content maintains the semantic integrity of the adverse event description, avoiding critical information truncation.
  • Use queries containing specific medical terminology and MedDRA codes. Verify if the recall results accurately return original text snippets containing these terms and check if Recall count meets information coverage requirements.

The values provided are common starting points and should be measured against the reader's own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.