Data Characteristics in this Category
Intelligent triage in pharmacovigilance focuses on patient medication information, medical history, allergy history, current symptom descriptions, drug inserts, and Adverse Drug Reaction (ADR) database content. Data sources include Electronic Health Record (EHR) systems, patient self-reports, pharmaceutical company drug information documents, and national ADR monitoring center data. Data updates are frequent, especially for symptom descriptions and ADR data, which may update in real-time or daily. Document structures are primarily semi-structured and unstructured, such as patient medical record text, drug insert PDFs, and free-text fields in ADR reports. Key fields include generic drug name, batch number, dosage, administration, patient age, gender, diagnosis, allergens, ADR event description, occurrence time, and severity. Units often include milligrams, milliliters, days, and times.
Constraints Imposed by these Features on Workflow Orchestration
High-frequency updates of patient symptom data and ADR reports require the workflow to support real-time or near real-time data ingestion and processing. This prevents outdated information from affecting triage accuracy. Semi-structured and unstructured text data, such as patient medical records and drug inserts, necessitate integrating advanced text parsing and entity recognition components within the workflow to accurately extract key information. The multi-source heterogeneous data characteristic requires the workflow to handle various data formats and field mappings during data integration, ensuring data consistency. For example, drug names may have aliases or abbreviations requiring standardization. The specialized nature of pharmacovigilance also dictates that the workflow must effectively process medical terminology and complex causal relationships, demanding high accuracy in model inference. Error handling mechanisms must be considered for missing or inconsistent data.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
UPLOAD_FILE_MAX_SIZE | 200 MB | Accommodates large drug inserts or ADR report files, ensuring smooth uploads. |
PARSE_FILE_TIMEOUT_SECONDS | 300 seconds | Provides sufficient time for parsing large PDFs or complex text files, preventing timeouts. |
Chunk size | 800 characters | Balances semantic completeness with model processing efficiency, avoiding information loss from overly long or short segments. |
Recall count | Top 8 entries | Recalls enough relevant drug information and ADR cases from the knowledge base to improve triage accuracy. |
Similarity threshold | 0.75 | Balances the relevance and breadth of recall results, preventing interference from irrelevant information while reducing false negatives. |
Max Concurrency | Calibrate by actual measurement | Adjust based on system resources and expected concurrent triage volume to ensure response speed and service stability. |
Three Common Pitfalls
- A workflow runtime error
offset 17typically indicates that the text parsing component failed to correctly recognize special characters or unusual formats in medical text, leading to an index offset. - A network error after file upload may result from
UPLOAD_FILE_MAX_SIZEbeing set too low, causing large files to exceed the upload limit. - A workflow stopping after two minutes often means
PARSE_FILE_TIMEOUT_SECONDSis insufficient for parsing complex medical documents, or model inference time is too long and not properly handled.
How to Verify Configuration
- Upload drug inserts and ADR reports in various formats (PDF, TXT, DOCX) and sizes. Observe if key fields, such as drug names and adverse reaction descriptions, are parsed and extracted correctly.
- Simulate patient symptom descriptions. Test if intelligent triage accurately links to relevant drugs and potential adverse reactions in the knowledge base. Verify if the number of recalled items matches expectations.
- Review workflow logs to confirm no
offseterrors or file parsing timeouts occur. Ensure smooth data flow between components. - Use medical terminology queries of varying complexity to test the effect of
Similarity threshold. Observe if the precision and coverage of recall results meet expectations.
The values provided are common starting points. Measure them against your own samples to determine the most suitable configuration.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.