Data Characteristics
Data in the pharmacovigilance domain for rational drug use primarily originates from Hospital Information Systems (HIS), Electronic Medical Records (EMR), pharmacy management systems, and national adverse drug reaction (ADR) reporting centers. Data updates are frequent; for inpatients, updates are typically daily, while for outpatients, updates occur after each visit. Document structures are diverse, including structured drug prescriptions, medication records, and patient diagnostic information, as well as unstructured physician progress notes, nurse observation records, and patient self-reports. Structured data fields include generic drug names, brand names, dosages, frequencies, administration routes, medication start/end times, patient age, gender, diagnoses, and allergy history. Unstructured text describes post-medication symptoms, signs, and abnormal laboratory findings. Units strictly adhere to international standards, such as mg, g, ml for drug dosages, d, h, min for time units, and various biochemical indicator units.
Constraints Imposed by Data Characteristics on Workflow Orchestration
The wide range of data sources requires workflows to support multi-source data ingestion and integration. High-frequency updates necessitate workflows that can support near real-time or scheduled triggers to ensure timely pharmacovigilance. The coexistence of structured and unstructured data challenges data preprocessing modules within the workflow. These modules must parse standard fields and extract key information from free text, for example, identifying potential adverse reaction events from physician progress notes. Strict unit specifications mean that unit standardization is required during data conversion and comparison to prevent incorrect judgments due to inconsistent units. The presence of sensitive patient information requires workflows to adhere to strict data security and privacy protection protocols during data processing and transmission, such as anonymizing patient identity information.
Configuration Guidelines
| Configuration Item | Suggested Value | Rationale |
|---|---|---|
maxContext | 8000 tokens | Balances long text processing with inference efficiency, ensuring the model can cover complete medication records and progress descriptions. |
Recall count | Top 20 entries | Increases recall rate for relevant medical knowledge and adverse reaction reports, covering various potential associations. |
Similarity threshold | 0.75 | Balances recall accuracy and relevance, avoiding the introduction of excessive irrelevant information that could affect judgment. |
Rerank result count | Top 5 entries | Selects the most relevant knowledge snippets for display, reducing cognitive load for subsequent judgment. |
Workflow Timeout Duration | 600 seconds | Accommodates complex data processing and multi-step inference, ensuring the workflow has sufficient time to complete tasks. |
API_KEY | Calibrate by actual measurement | Ensures access permissions for external knowledge bases or service calls, guaranteeing data security and access control. |
Common Pitfalls
- When processing unstructured progress notes, the model may fail to identify all potential adverse reaction symptoms, leading to underreporting. This occurs because keyword or pattern matching rules in the text extraction phase are not comprehensive enough to cover various expressions of medical terminology.
- The workflow may receive a
403 Forbiddenerror code when calling an external drug knowledge base, causing drug information matching to fail. This is typically due to an incorrect or expiredAPI_KEYconfiguration, preventing authentication with the external service. - The workflow may encounter a
workflow_execution_timeoutprompt when processing a large volume of patient data, resulting in incomplete vigilance analysis for some patients. This happens when theWorkflow Timeout Duration(workflow timeout) is set too short, unable to handle scenarios with large data volumes or complex inference steps.
Verification Steps
- Select typical medication cases, including scenarios with known adverse reactions and no adverse reactions. Perform end-to-end testing through the workflow to check if the output accurately identifies adverse reaction signals.
- Verify that when processing structured data, key fields such as drug names, dosages, and frequencies are correctly parsed and mapped to the internal data model.
- Examine workflow logs to confirm that all external API calls (e.g., drug knowledge base, adverse reaction database) return a
200 OKstatus code and that data transfer is normal. - Compare the adverse reaction reports processed by the workflow with expert manual judgments. Evaluate consistency and adjust the
Similarity threshold(similarity threshold) and model parameters based on feedback from medical experts.
Note: The values provided are common starting points and should be measured against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.