Workflow Orchestration for Rare Disease Pharmacovigilance

Rare disease pharmacovigilance data comes from diverse sources. These include global drug regulatory adverse event reporting systems (e.g., FDA

Data Characteristics in This Category

Rare disease pharmacovigilance data comes from diverse sources. These include global drug regulatory adverse event reporting systems (e.g., FDA Adverse Event Reporting System, FAERS), medical literature, clinical trial data, patient registries, and social media. Data update frequencies vary. Regulatory reports typically release quarterly or annually. Medical literature and clinical trial data generate continuously. Document structures are complex. They often contain unstructured text descriptions, semi-structured medical terminology (e.g., MedDRA codes), and structured patient demographic information. Fields cover patient demographics, medication history, adverse event descriptions, event outcomes, and diagnostic information. Reports often involve multiple languages. They also lack uniform unit standards. For example, dosages may be expressed in milligrams, micrograms, or international units. Time units may mix hours, days, and months.

Constraints Imposed by These Characteristics on "Workflow Orchestration"

Irregular rare disease data updates require flexible workflow triggers. These triggers must adapt to a combination of periodic pulling and real-time event-driven patterns. Diverse and heterogeneous data structures, especially large amounts of unstructured text, demand high natural language processing (NLP) capabilities during data preprocessing. This requires efficient entity recognition, relationship extraction, and event classification models. The lack of uniform unit standards requires workflows to integrate unit conversion modules during the data cleaning phase. This ensures comparability of numerical fields. Additionally, the small number of rare disease patients results in fewer adverse event reports. Workflow design must consider model robustness in small-sample learning scenarios. It must also effectively use external knowledge, such as knowledge graphs, for auxiliary judgment. Multilingual reports also demand language identification and translation capabilities from text processing modules. This ensures information completeness and accuracy.

Configuration Settings

Configuration ItemRecommended ValueRationale
maxContext2000–3000 tokensRare disease adverse event reports often contain detailed medical history and event descriptions. A longer context window is needed to capture key information.
Chunk size (Segment Length)500 characters (characters)For unstructured text, an appropriate segment length helps maintain semantic integrity and prevents truncation of critical information.
Recall count (Recall Count)Top 10 entries (top 10 items)Rare disease knowledge base entries might be limited. Increasing the recall count can improve coverage of relevant information.
Similarity threshold (Similarity Threshold)0.78Rare disease symptoms and adverse reaction descriptions may have subtle differences. A higher threshold helps precise matching and avoids interference from irrelevant information.
Rerank result count (Rerank Return Count)Top 5 entries (top 5 items)Reranking based on recall focuses on the most relevant few items, improving large model processing efficiency.
PARSE_FILE_TIMEOUT_SECONDS600 seconds (seconds)Rare disease report documents may contain complex tables or lengthy descriptions. Allowing sufficient time for parsing prevents timeouts.

Three Common Mistakes

  • The large model hallucinates when processing adverse event reports. Its replies do not align with facts cited from the knowledge base. This happens because the workflow lacks a fact-checking mechanism for large model generated content. It fails to compare and verify the original snippets cited from the knowledge base against the model's output.
  • Specific drug dosages or laboratory indicators are not correctly identified or converted. This leads to deviations in subsequent analysis results. This occurs because the unit standardization module in the data preprocessing stage does not cover all rare disease-specific or non-standardized measurement units.
  • Workflow execution times out or fails. Logs show file parsing errors. This happens because processed PDF documents contain scanned images or complex charts. Default text extractors cannot effectively parse this non-text content.

How to Confirm Proper Configuration

  • Select a test set containing typical adverse event reports. Run the workflow and check the accuracy of key entity recognition (e.g., drugs, adverse events, dosages) in the final output. Ensure it meets the predefined minimum accuracy threshold.
  • Verify the workflow's unit conversion function. Input dosage or time data with different units. Check if the output has been standardized to uniform units. Compare results with manual verification to confirm accurate conversion.
  • Check the effectiveness of the knowledge base citation verification mechanism. Submit a question known to cause hallucination. Observe whether the large model identifies unreasonable citations and flags or corrects them. Cross-reference the flagged unreasonable citations with the actual knowledge base source text.

The values provided are common starting points. Measure them against your own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.