Data Characteristics in this Category
E-pharmacy platforms generate pharmacovigilance data primarily from user-submitted adverse event reports, product reviews, consultation records, and drug sales and logistics data. This data often exists as unstructured text, semi-structured JSON, or structured CSV. Adverse event reports have a high update frequency; initial classification and processing are typically required within minutes of user submission. Product reviews and consultation records are generated in real-time based on user behavior.
Regarding document structure, adverse event reports commonly include fields such as patient basic information, medication history, adverse event description (symptoms, onset time, duration, outcome), and suspected drug information (name, batch number, dosage, administration). Field content is often natural language descriptions, involving medical terminology, colloquialisms, and even typos. Units frequently include milligrams (mg), grams (g), milliliters (mL), and tablets for dosage, and times/day or week for frequency.
Constraints Imposed by These Characteristics on Workflow Orchestration
The data characteristics of e-pharmacy impose specific requirements on workflow orchestration. The real-time nature of adverse event reports demands rapid workflow response and support for high-concurrency processing. For example, setting maxContext to a higher value can prevent queue buildup. The prevalence of unstructured text data necessitates a strong reliance on Natural Language Processing (NLP) nodes for information extraction and entity recognition, such as using JsonPath to extract key medical entities from unstructured text. The presence of colloquialisms and typos requires NLP models to be robust and may necessitate the introduction of correction or standardization steps. The diversity of fields in structured data requires flexible mapping capabilities during data cleaning and transformation within the workflow. Furthermore, the presence of sensitive information like drug batch numbers demands high requirements for data anonymization and access control nodes to ensure compliance during data flow.
Configuration Guidelines
| Configuration Item | Suggested Value | Rationale |
|---|---|---|
maxContext | 50 | Addresses high-concurrency adverse event report submissions, ensuring rapid response. |
Chunk size (Chunk Length) | 800-1200 characters | Balances the detail level of adverse event reports with model processing efficiency, reducing truncation. |
Recall count (Recall Count) | Top 10 | Ensures coverage of relevant background knowledge (e.g., drug inserts, historical similar reports) to aid judgment. |
Similarity threshold (Similarity Threshold) | 0.75 | Filters for historical data highly relevant to the current adverse event report, improving classification accuracy. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Allows sufficient time to process complex drug inserts or unstructured data within attachments. |
Rerank result count (Rerank Return Count) | Top 3 | Focuses on the most critical information, reducing manual screening costs for engineers. |
Three Common Mistakes
- Incorrect
JsonPathexpression when extracting JSON content fromhttp responsein an HTTP request node, leading to downstream nodes receiving null or incorrect data. This occurs due to complex JSON structures or expressions failing to accurately match target fields. base_urlis not configured correctly in a daily report workflow, preventing the AI model from accessing external knowledge bases or API services. This happens when the actual deployment address and port of the API service are not carefully verified.- A workflow exported to another environment fails to run, indicating missing Python dependency files. This occurs when not all custom scripts or environment configurations are included during export, leading to a lack of necessary runtime components in the target environment.
How to Confirm Correct Configuration
- Submit a simulated adverse event report and observe the workflow logs to confirm that all nodes execute successfully and that key information (e.g., drug name, adverse event type) from the
http responseis correctly extracted. - Add a test step to the workflow to check if the external API configured in
base_urlcan successfully return data, for example, via a ping or a simple GET request. - Export the workflow and import it into a new environment to run, confirming that all custom Python scripts and background knowledge are correctly loaded and called, with no missing files or path errors.
- Use a set of test data with known classification results, process it through the workflow, and compare the final output with the expected classification to confirm that parameters like
Similarity threshold(Similarity Threshold) andRecall count(Recall Count) are appropriately set.
The values given are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.