Deployment and Upgrade for CAR-T Cell Therapy Pharmacovigilance

CAR-T cell therapy pharmacovigilance data is high-dimensional and highly heterogeneous. Data sources include clinical trial reports, real-world

Data Characteristics

CAR-T cell therapy pharmacovigilance data is high-dimensional and highly heterogeneous. Data sources include clinical trial reports, real-world evidence (RWE) data, electronic health records (EHR), patient-reported outcomes (PROs), and post-market surveillance reports submitted by regulatory agencies. Data update frequencies vary; clinical trial data updates typically occur with interim study reports, while post-market surveillance data may be submitted quarterly or annually. Document structures often combine unstructured text (e.g., medical narratives, adverse event descriptions) and semi-structured data (e.g., tabular laboratory indicators, diagnostic codes). Fields include common adverse event (AE) types, severity, and occurrence time, as well as specific fields such as cytokine release syndrome (CRS) grading, immune effector cell-associated neurotoxicity syndrome (ICANS) scores, CAR-T product batch numbers, and lymphocyte depletion regimens. Units involve biomarker indicators like cell counts (e.g., cells/μL) and cytokine concentrations (e.g., pg/mL).

Constraints on Deployment and Upgrade

The heterogeneity and high dimensionality of CAR-T cell therapy pharmacovigilance data impose specific requirements on FastGPT deployment and upgrades. Diverse data sources necessitate robust data ingestion capabilities, handling various document formats and database connections. Inconsistent update frequencies require the knowledge base update mechanism to support flexible configuration of incremental updates and periodic full re-indexing, avoiding unnecessary resource consumption. The high proportion of unstructured text dictates a reliance on natural language processing (NLP) capabilities, particularly deep understanding of medical terminology and clinical context. The presence of specific fields, such as CRS grading and ICANS scores, requires customized entity recognition and information extraction models to ensure accurate capture of critical information. These characteristics collectively constrain FastGPT's practices in data preprocessing, model selection, knowledge base construction strategies, and resource configuration to meet the specialized needs of CAR-T pharmacovigilance.

Configuration Guidelines

Configuration ItemRecommended ValueRationale
UPLOAD_FILE_MAX_SIZE200 MBClinical trial reports and real-world study documents often contain numerous charts, graphs, and detailed text, resulting in large file sizes.
maxContext3000 TokensAccommodates long text content in medical narratives and adverse event descriptions, ensuring context completeness.
PARSE_FILE_TIMEOUT_SECONDS600Large file parsing and complex document structures require longer timeout periods to prevent parsing failures.
Chunk size (Chunk Length)500–800 characters (characters)Balances the completeness of medical concepts with the granularity of retrieval, preventing critical information truncation.
Similarity threshold (Similarity Threshold)0.75Ensures retrieval of knowledge segments highly relevant to CAR-T specific adverse events, reducing noise.
Rerank result count (Reranked Return Count)Top 5 entries (top 5)Further prioritizes key information through reranking based on high-similarity retrieval.

Common Pitfalls

  • Observation: After a knowledge base update, some adverse event descriptions are not correctly recognized or linked. Reason: Lack of customized entity recognition rules for CAR-T specific terminology (e.g., neurotoxicity syndrome, cytokine storm).
  • Observation: Data import tasks frequently time out, especially when processing large clinical trial reports. Reason: PARSE_FILE_TIMEOUT_SECONDS is set too low, insufficient for handling PDF files with extensive unstructured text and complex tables.
  • Observation: In a private deployment environment, AI model inference speed is significantly lower than expected, leading to user response delays. Reason: The deployment environment lacks sufficient GPU resource configuration, or large embedding models like bge-m3 do not fully utilize GPU acceleration, causing computational bottlenecks.

Validation Steps

  • Select at least 5 representative CAR-T clinical trial reports and real-world study documents. Execute the full data import process. Check if all key fields (e.g., CRS grading, ICANS score, CAR-T product batch) are accurately extracted and stored.
  • Simulate more than 10 queries involving CAR-T specific adverse events. Evaluate the proportion of knowledge segments containing specific terminology in the retrieval results. This proportion should be significantly higher than for general pharmacovigilance queries.
  • Monitor the execution logs of knowledge base update tasks. Confirm that all data source types (e.g., PDF, JSON, database records) complete processing within the preset PARSE_FILE_TIMEOUT_SECONDS without timeout errors.
  • Check system logs for error messages related to FastGPT containers or bge-m3 model services to ensure all components are operating correctly.

The values provided are common starting points and should be measured against specific samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.