Data Characteristics
CAR-T cell therapy pharmacovigilance data originates from clinical trial reports, real-world evidence (RWE) studies, patient follow-up records, adverse drug reaction (ADR) reporting systems, and literature databases. This data updates frequently; post-market surveillance ADR reports can update daily. Document structures vary, including unstructured free text (e.g., physician notes, patient narratives), semi-structured case report forms (CRF), and structured database entries. Fields cover patient demographics, treatment plans, medical history, complications, ADR types, severity, onset time, duration, interventions, and outcomes. ADR names typically follow MedDRA coding. Dosage units involve cell counts like cells/kg or cells/m^2.
Constraints Imposed by These Characteristics on Workflow Orchestration
The heterogeneous nature of CAR-T cell therapy pharmacovigilance data requires robust data integration and cleaning capabilities in the workflow. High-frequency updates mean the workflow must support real-time or near real-time triggering and processing to ensure timely identification and evaluation of adverse events. The presence of unstructured text necessitates integrating natural language processing (NLP) modules to accurately extract key information from vast medical texts, such as ADR symptoms, onset times, and drug associations. The need for MedDRA coding standardization constrains the workflow to include medical terminology mapping in its data standardization nodes. Furthermore, CAR-T therapy's specific cell count units require dedicated unit conversion or range check logic during data validation.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Data Source Connector | FHIR, CDISC ODM, Custom API | Integrate clinical data and ADR reporting systems from diverse sources |
NLP_Model Selection | Pre-trained medical NLP models, e.g., BioBERT | Accurately identify medical terms, entities, and relationships in unstructured text |
Chunk size | 500–800 characters | Balance semantic completeness with model processing efficiency, reducing context loss |
Recall count | Top 10–15 entries | Ensure retrieval of sufficient relevant information to cover potential ADR associations |
Similarity threshold | 0.75–0.85 | Precisely match similar adverse event incidents, reducing false positives |
MedDRA_Encoding Mapping | ICD-10 to MedDRA dictionary mapping | Standardize adverse event terminology for statistical analysis and regulatory reporting |
Three Common Pitfalls
- Symptom: Workflow calls an external knowledge base, but the returned result is empty or incomplete. Cause: External knowledge base API parameters are incorrectly mapped, or the query statement format does not meet the target system's requirements.
- Symptom: When processing a large volume of patient follow-up records, the workflow experiences timeout or out-of-memory errors. Cause: Data batch processing logic is missing, or a single processing unit handles too much data, exceeding system resource limits.
- Symptom: When passing the output of one workflow to another, the downstream workflow fails to execute after the
User Selectionnode. Cause: The upstream workflow's output data format does not match the input expected by the downstream workflow'sUser Selectionnode, preventing the node from parsing or triggering correctly.
Validation Steps
- Simulate real ADR reports to verify if the workflow accurately identifies adverse reactions and extracts key information.
- Check workflow logs to confirm all data source connectors establish successful connections and data synchronization occurs without errors.
- Perform manual spot checks on NLP-processed data to assess the accuracy of medical term extraction and MedDRA code mapping.
- Run test cases with large data volumes to monitor workflow execution time and resource consumption, ensuring system stability.
Note: The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.