Workflow Orchestration for Solid Tumor Pharmacovigilance

Solid tumor pharmacovigilance data originates from clinical trial reports, real-world evidence (RWE) data, case reports, medical literature, and

Data Characteristics in This Category

Solid tumor pharmacovigilance data originates from clinical trial reports, real-world evidence (RWE) data, case reports, medical literature, and adverse drug reaction (ADR) reporting systems. Data update frequencies vary. Clinical trial data typically updates after phased study results are released, while ADR reporting systems may update daily or weekly. Document structures are diverse. They include unstructured free-text descriptions, semi-structured tabular data (e.g., patient demographics, medical history, adverse event descriptions, laboratory test results), and structured coded data (e.g., MedDRA codes). Fields include general demographic information, tumor type (e.g., ICD-O-3 codes), tumor stage, pathological features, genetic mutation information (e.g., EGFR, ALK mutation status), treatment regimens (drug name, dosage, administration route, cycle), adverse event severity (CTCAE grades), and outcomes. Units for dosage commonly use milligrams (mg), grams (g), or milligrams per kilogram (mg/kg). Time units often use days (day), weeks (week), or months (month). Laboratory indicators use their own international standard units based on the specific test item.

Constraints Imposed by These Characteristics on Workflow Orchestration

The diversity of solid tumor pharmacovigilance data imposes specific requirements on workflow orchestration. First, the high proportion of unstructured text, such as detailed adverse event descriptions from clinicians, requires robust Natural Language Processing (NLP) capabilities for information extraction and entity recognition. This necessitates integrating text parsing and entity linking tools into the workflow, configured with appropriate dictionaries and rules. Second, varying data source update frequencies require flexible trigger mechanisms in the workflow. For example, scheduled fetching can be set for ADR reporting systems, while literature data can be triggered on demand. The complexity of document structures, especially semi-structured tables, demands that the workflow can handle data import in different formats, along with field mapping and standardization. For solid tumor-specific fields like tumor staging and genetic mutation information, knowledge base matching within the workflow requires the knowledge base to include these specialized terms and their synonyms. Adverse event severity grading, such as CTCAE grades, influences risk assessment logic. The workflow needs to perform conditional judgments and branching based on these grades. Finally, the coexistence of multiple units for laboratory indicators requires unit standardization or conversion during data processing to avoid calculation errors.

Configuration Guidelines

Configuration ItemSuggested ValueRationale
Chunk size800–1200 charactersSolid tumor case reports often contain detailed descriptions; balances context and processing efficiency
Recall countTop 8 entriesEnsures coverage of knowledge points related to adverse events, drugs, and patient characteristics
Similarity threshold0.75–0.85Solid tumor terminology is highly specific; a higher threshold reduces false positives
PARSE_FILE_TIMEOUT_SECONDS600 secondsProcessing large clinical trial reports or multi-page PDF documents may require more time
maxContext4000 charactersEnsures accommodation of complete adverse event descriptions and relevant background information
Rerank result countTop 3 entriesFocuses on the most relevant knowledge points, reducing subsequent processing burden

Three Common Mistakes

  • Tool call results do not output as expected: This may be due to a missing explicit output node in the workflow, or the output node is configured with an incorrect output target.
  • Document parsing node reports a 404 error for files on the server: This typically results from server-side permission issues, preventing FastGPT from accessing or reading the uploaded file path, or incorrect file path mapping.
  • Workflow fails to accurately filter based on solid tumor-specific fields (e.g., 基因突变): This often occurs when the knowledge base lacks corresponding specialized vocabulary or synonyms, leading to vector matching failure, or when the conditional judgment logic in the workflow does not correctly reference these fields.

How to Confirm Correct Configuration

  • Select test cases containing typical solid tumor adverse events. Run the workflow and check if the final output accurately extracts key information such as drugs, adverse events, CTCAE grades, and relevant genetic mutations.
  • Upload solid tumor-related documents in different formats (PDF, DOCX, TXT) and lengths (short case reports, long clinical trial reports). Observe if the document parsing node successfully processes all of them and check logs for timeouts or errors.
  • Add or modify some specialized solid tumor terms and their synonyms in the knowledge base. Rerun relevant test cases to verify if the workflow's recall and judgment logic correctly respond to these updates.

The values provided are common starting points. Measure them against your own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.