Data Characteristics in This Domain
Medical record quality control in pharmacovigilance primarily uses data from Electronic Health Record (EHR) systems. This data typically combines unstructured text, semi-structured tables, and structured fields. Data generation is real-time or near real-time as patients visit and receive treatment. However, quality control analysis often occurs periodically, such as daily, weekly, or monthly batch processing. Document structures are complex, including chief complaints, history of present illness, past medical history, medication records, examination and test results, diagnoses, and treatment plans. Fields and units involve drug names, dosages, frequencies, administration routes, medication times, adverse event descriptions, vital signs, and laboratory indicators with their units. Drug names may have aliases and abbreviations, dosage units vary (e.g., mg, g, IU, ml), and adverse event descriptions are often natural language text.
Constraints Imposed by These Characteristics on "Deployment and Upgrade"
The mixed structure of medical record data and its unstructured text characteristics require robust text parsing and entity recognition capabilities for knowledge base construction to accurately extract drug information and adverse events. Diverse data sources necessitate support for multiple data interfaces and formats during data ingestion, followed by standardization. Periodic quality control analysis requires deployment solutions that include batch processing capabilities and scheduled task orchestration. The variety of drug aliases and dosage units demands a comprehensive dictionary and ontology for the knowledge base, requiring maintenance of extensive synonyms and unit conversion rules. The natural language nature of adverse event descriptions emphasizes the performance of Retrieval-Augmented Generation (RAG) systems in semantic understanding and precise recall to avoid underreporting or misreporting potential adverse drug reactions.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
UPLOAD_FILE_MAX_SIZE | 200 MB | Electronic medical record files are often large and may include images or attachments, requiring sufficient upload capacity. |
maxContext | 3000 characters | Medical record text paragraphs are long, requiring a larger context window to maintain semantic integrity and ensure drug information and adverse event descriptions are not truncated. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Parsing large medical record files is time-consuming, requiring a longer timeout to prevent parsing interruptions. |
Chunk size | 800–1200 characters | Extend segment length appropriately to capture associations between drugs and adverse events without splitting critical information. |
Recall count | Top 8 entries | Pharmacovigilance demands high accuracy and comprehensiveness in recall; increasing the number of recalled items covers more potential associations. |
Similarity threshold | Calibrate by actual measurement | Pharmacovigilance requires high precision in recall to avoid interference from irrelevant information; fine-tuning on test sets is necessary. |
Rerank result count | Top 5 entries | Re-ranking further improves the order of critical information based on recall, ensuring the most relevant adverse event leads are prioritized. |
Three Common Pitfalls
- After building a Docker image, an error message appears about an unconfigured commercial link or a
getPluginGroups500 error. The system may fail to start or plugin functionality may be unavailable. This usually stems from configuration differences between open-source and commercial versions, or incorrect loading of environment variables during Docker build. - Uploading large medical record files results in timeouts or upload failures, with the file upload progress bar stalling or a
Connection timed outerror. This likely occurs becauseUPLOAD_FILE_MAX_SIZEorPARSE_FILE_TIMEOUT_SECONDSare set too low for the characteristics of medical record files. - Drug names or dosage information are missing or incorrect in RAG results. Generated summaries or analysis reports may have empty or garbled key fields. This can be due to the knowledge base's entity recognition model having insufficient capability to identify diverse drug aliases and units in medical records, or a segmentation strategy that fragments critical information.
How to Confirm Correct Configuration
- Upload a real medical record file containing typical adverse drug reaction descriptions. Check if the file is successfully parsed and segmented. Verify that core entities like drug names, dosages, and adverse events are correctly extracted into the knowledge base.
- Perform a search for specific drug and adverse event combinations. Observe if the recalled medical record snippets in the RAG results are accurately associated. Check if the generated summaries effectively reflect pharmacovigilance-related information. Evaluate if recall and re-ranking performance meet business expectations.
- Conduct batch medical record data import and parsing tests. Monitor system resource utilization to confirm stable operation without memory leaks or CPU overload under expected data volumes and update frequencies. Ensure processing speed meets quality control cycle requirements.
Note: The values provided are common starting points. They should be measured against specific samples and use cases.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.