Workflow Orchestration for CAR-T Cell Therapy Pharmacovigilance

CAR-T cell therapy pharmacovigilance data originates from clinical trial reports, real-world evidence (RWE) studies, patient follow-up records

Data Characteristics

CAR-T cell therapy pharmacovigilance data originates from clinical trial reports, real-world evidence (RWE) studies, patient follow-up records, adverse drug reaction (ADR) reporting systems, and literature databases. This data updates frequently; post-market surveillance ADR reports can update daily. Document structures vary, including unstructured free text (e.g., physician notes, patient narratives), semi-structured case report forms (CRF), and structured database entries. Fields cover patient demographics, treatment plans, medical history, complications, ADR types, severity, onset time, duration, interventions, and outcomes. ADR names typically follow MedDRA coding. Dosage units involve cell counts like cells/kg or cells/m^2.

Constraints Imposed by These Characteristics on Workflow Orchestration

The heterogeneous nature of CAR-T cell therapy pharmacovigilance data requires robust data integration and cleaning capabilities in the workflow. High-frequency updates mean the workflow must support real-time or near real-time triggering and processing to ensure timely identification and evaluation of adverse events. The presence of unstructured text necessitates integrating natural language processing (NLP) modules to accurately extract key information from vast medical texts, such as ADR symptoms, onset times, and drug associations. The need for MedDRA coding standardization constrains the workflow to include medical terminology mapping in its data standardization nodes. Furthermore, CAR-T therapy's specific cell count units require dedicated unit conversion or range check logic during data validation.

Configuration Guidelines

Configuration ItemRecommended ValueRationale
Data Source ConnectorFHIR, CDISC ODM, Custom APIIntegrate clinical data and ADR reporting systems from diverse sources
NLP_Model SelectionPre-trained medical NLP models, e.g., BioBERTAccurately identify medical terms, entities, and relationships in unstructured text
Chunk size500–800 charactersBalance semantic completeness with model processing efficiency, reducing context loss
Recall countTop 10–15 entriesEnsure retrieval of sufficient relevant information to cover potential ADR associations
Similarity threshold0.75–0.85Precisely match similar adverse event incidents, reducing false positives
MedDRA_Encoding MappingICD-10 to MedDRA dictionary mappingStandardize adverse event terminology for statistical analysis and regulatory reporting

Three Common Pitfalls

  • Symptom: Workflow calls an external knowledge base, but the returned result is empty or incomplete. Cause: External knowledge base API parameters are incorrectly mapped, or the query statement format does not meet the target system's requirements.
  • Symptom: When processing a large volume of patient follow-up records, the workflow experiences timeout or out-of-memory errors. Cause: Data batch processing logic is missing, or a single processing unit handles too much data, exceeding system resource limits.
  • Symptom: When passing the output of one workflow to another, the downstream workflow fails to execute after the User Selection node. Cause: The upstream workflow's output data format does not match the input expected by the downstream workflow's User Selection node, preventing the node from parsing or triggering correctly.

Validation Steps

  • Simulate real ADR reports to verify if the workflow accurately identifies adverse reactions and extracts key information.
  • Check workflow logs to confirm all data source connectors establish successful connections and data synchronization occurs without errors.
  • Perform manual spot checks on NLP-processed data to assess the accuracy of medical term extraction and MedDRA code mapping.
  • Run test cases with large data volumes to monitor workflow execution time and resource consumption, ensuring system stability.

Note: The values provided are common starting points and should be measured against the reader's own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.