Data Characteristics
Gene therapy AAV (adeno-associated virus vector) pharmacovigilance data originates from clinical trial reports, real-world studies, post-market surveillance, and global adverse event reporting systems. This data is highly heterogeneous, typically including detailed patient demographics, gene therapy product batch information, administration routes, dosages, adverse event (AE) descriptions, severity, occurrence time, outcomes, causality assessments, and interventions. Data update frequency is not fixed; it is usually quarterly or annually during clinical trials, and real-time or periodic aggregation occurs post-market based on reports. Document structure is complex, potentially containing semi-structured medical text, structured laboratory indicators, imaging reports, and pathology reports. Fields and units vary; for example, laboratory indicators use both SI and conventional units, while adverse event descriptions are often free text requiring standardized encoding (e.g., MedDRA coding).
Constraints on Workflow Orchestration
The complexity of AAV gene therapy pharmacovigilance data imposes specific requirements on workflow orchestration. Data diversity necessitates workflow support for multi-source heterogeneous data ingestion and preprocessing modules. Free-text adverse event descriptions require robust text processing capabilities, including entity recognition and event extraction, which mandates integrating advanced Natural Language Processing (NLP) components into the workflow. Varying update frequencies between clinical trial data and post-market surveillance data mean the workflow should flexibly configure trigger mechanisms, supporting periodic batch processing and real-time report responsiveness. The use of specialized terminology like MedDRA coding requires a knowledge base with professional dictionaries and mapping during the semantic understanding phase of the workflow. Furthermore, adverse event severity assessment and causality determination involve complex logic, requiring sophisticated validators and decision branches within the workflow to guide subsequent processing paths, such as triggering expert review or automatic report generation. Processing this data requires high fault tolerance, handling missing values, inconsistent data, and retry mechanisms for failures.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
maxContext | 8000 tokens | AAV adverse event reports often contain multiple detailed descriptions and medical terms, requiring a long context window for complete semantic understanding. |
Chunk size (Segment Length) | 500 characters (characters) | Ensures each text segment contains sufficient information for semantic analysis while avoiding excessive length that could lead to information redundancy or reduced processing efficiency. |
Recall count (Recall Count) | Top 10 entries (top 10) | Gene therapy adverse event knowledge points may be distributed across different documents; increasing recall appropriately helps improve relevance coverage. |
Similarity threshold (Similarity Threshold) | 0.75 | Given the specialized and rigorous nature of medical text, a higher similarity threshold reduces the recall of irrelevant or low-quality information. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds (seconds) | AAV pharmacovigilance documents often include numerous charts and complex layouts, and parsing can be time-consuming; ample time prevents timeout interruptions. |
Rerank result count (Reranked Return Count) | Top 5 entries (top 5) | After high recall, reranking selects a smaller number of the most relevant pieces of information, improving the precision and conciseness of the final output. |
Common Pitfalls
- Symptom: The workflow ends immediately after knowledge base retrieval without proceeding to subsequent judgment or report generation steps. Cause: The output of the knowledge base retrieval node is not correctly connected to subsequent nodes, or the
input parametersof subsequent nodes are misconfigured, preventing the process from continuing. - Symptom: The pharmacovigilance report generated by the large language model (LLM) has inconsistent formatting or lacks critical information. Cause: The LLM prompt does not clearly constrain the output format and required key fields, or the
custom typevariable is not fully utilized to dynamically inject necessary context into the prompt. - Symptom: When processing reports containing MedDRA codes, the workflow fails to correctly identify or map the codes. Cause: The latest MedDRA dictionary is not imported into the knowledge base, or corresponding encoding mapping rules are not configured in the entity recognition module of the workflow.
Verification Steps
- Select a test set containing typical AAV adverse event reports. Run the workflow and check if the final generated report includes all key information. Compare it against the expected output to ensure information completeness.
- Set up log outputs at key nodes of the workflow (e.g., knowledge base retrieval, LLM processing, validators). Observe intermediate results in the logs to confirm data flow and processing logic align with the design.
- Test the workflow's decision branches for adverse event cases with different severities and causalities. Verify if it correctly guides to the appropriate processing paths, such as triggering expert review notifications.
- Simulate abnormal data (e.g., missing key fields, malformed reports). Run the workflow and check if error handling mechanisms effectively capture issues and provide prompts, or if they perform retries according to predefined logic.
Note: The values given are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.