Data Characteristics
CAR-T cell therapy pharmacovigilance data primarily originates from clinical trial reports, real-world evidence (RWE) data, case reports, post-market surveillance reports, and regulatory guidelines. These documents are typically in PDF format, with some in Word or exported from structured databases. Data updates frequently, especially during new product launches and regulatory policy changes. Document structures are complex, containing extensive medical terminology, abbreviations, charts, and tables. Information includes patient demographics, CAR-T product details, adverse event (AE) descriptions, serious adverse event (SAE) reports, laboratory test results, concomitant medications, and treatment outcomes. Fields may include standardized medical codes (e.g., MedDRA codes) and free-text descriptions. Units involve dosage (e.g., cells/kg), time (e.g., days, weeks), and laboratory indicators (e.g., pg/mL, %). Field names and unit representations can vary across different source documents.
Constraints on Document Parsing and Chunking
The complex structure and high information density of CAR-T cell therapy documents demand advanced parsing capabilities. Large volumes of charts and tables require accurate extraction of key data, preventing information loss. The prevalence of medical terminology and abbreviations necessitates preserving contextual integrity during chunking to avoid semantic ambiguity caused by truncation. For example, a complete adverse event description might span multiple paragraphs or include tabular data. Mechanically chunking by paragraph or fixed character count can fragment critical information. High data update frequency requires the parsing and chunking process to support efficient incremental updates and version management, ensuring knowledge base timeliness. Diverse fields and units require parsed content to retain its original context during chunking, facilitating subsequent semantic understanding and retrieval. For clinical trial reports often hundreds of pages long, parser stability and resource consumption are key considerations.
Configuration Recommendations
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
UPLOAD_FILE_MAX_SIZE | 500 MB | Clinical trial reports can be large; ensure full upload capability. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Large file parsing is time-consuming; prevent timeout failures. |
Chunk size | 800–1200 characters | Balances contextual completeness and retrieval efficiency; avoids overly long or short chunks. |
Chunk Overlap Length | 100–200 characters | Ensures semantic continuity at chunk boundaries, especially for medical terms and descriptions. |
maxContext | 3 | For complex medical texts like CAR-T, retaining more context aids understanding. |
Recall count | Top 5-8 entries | Increases recall rate of relevant information, covering different angles of adverse event descriptions. |
Common Pitfalls
- Parsing hundreds of PDF pages fails, while parsing tens of pages succeeds. This typically occurs when
PARSE_FILE_TIMEOUT_SECONDSis set too low, causing the parser to time out when processing large, complex documents. - Key laboratory indicator values or dosage information are missing from chunked recall results. This usually happens when
Chunk sizeis set too small, leading to the fragmentation of tables or chart captions containing critical values, preventing the formation of complete semantic blocks. - Locally deployed
pdf-markerthrows aCannot read properties of undefinederror. This issue may relate to compatibility between thepdf-markerversion and the FastGPT version, or incorrect environment dependency installation.
Verification Steps
- Upload a CAR-T clinical trial report containing complex tables and charts. Check if the parsed text content fully retains table data and chart descriptions.
- For a typical adverse event report document, adjust
Chunk sizeandChunk Overlap Length. Then, use the knowledge base preview function to verify if the chunked content maintains the semantic integrity of the adverse event description, avoiding critical information truncation. - Use queries containing specific medical terminology and MedDRA codes. Verify if the recall results accurately return original text snippets containing these terms and check if
Recall countmeets information coverage requirements.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.