Data Characteristics for This Category
Pharmacovigilance data in hospital operations primarily originates from the hospital's electronic medical record (EMR) system, pharmacy management system, and adverse event reports manually entered by clinical staff. This data typically exists as unstructured text (e.g., doctor's orders, progress notes, nurse's notes) and semi-structured data (e.g., adverse reaction report forms). Data updates are frequent; adverse event reports can occur at any time, and EMR data is written in real-time. Document structures are diverse, including free-text descriptions, structured fields (e.g., drug name, dosage, administration route, adverse reaction type, occurrence time, severity assessment), and medical terminology codes (e.g., ICD-10, SNOMED CT). Fields and units involve drug generic names, brand names, batch numbers, dosage units (mg, g, ml, IU, etc.), timestamps, patient IDs, and department codes.
Constraints Imposed by These Characteristics on "Deployment and Upgrades"
The high real-time nature of hospital operations pharmacovigilance data requires FastGPT's knowledge base to synchronize updates quickly. This ensures the AI assistant provides information based on the latest data. Diverse document structures demand advanced data preprocessing and embedding models capable of supporting text extraction and semantic understanding across various formats. Medical terminology and abbreviations in free text, in particular, require specialized dictionaries or domain-specific model support. The large and continuously growing data volume challenges storage capacity and indexing efficiency. Furthermore, given patient privacy and sensitive medical information, private deployment is a priority, with strict requirements for data security and access control. Post-deployment upgrades must ensure business continuity and minimize service downtime.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
UPLOAD_FILE_MAX_SIZE | 500 MB | Accommodates potentially large individual medical record files or batch reports, ensuring complete uploads. |
maxContext | 1000 Tokens | Balances context length for lengthy progress notes, improving comprehension coherence. |
PARSE_FILE_TIMEOUT_SECONDS | 300 seconds | Allows sufficient time to process complex or large unstructured texts, such as tens of thousands of words in progress notes. |
Chunk size | 800–1200 characters | Adapts to the semantic density of medical texts, ensuring each segment contains enough information for retrieval. |
Recall count | Top 8 entries | Increases coverage when retrieving relevant adverse reaction cases or drug information from a massive knowledge base. |
Similarity threshold | 0.75 | Filters out irrelevant or weakly related retrieval results, improving accuracy and reducing false positives. |
Three Common Mistakes
- Knowledge base data is not updated promptly, leading to AI responses based on outdated information. This occurs because the data synchronization mechanism is not configured or executed frequently enough.
- The AI fails to understand medical professional terms or abbreviations, resulting in inaccurate or incomplete Q&A results. This happens when the embedding model is not optimized for the medical domain or lacks professional dictionaries.
- Service experiences prolonged interruptions during deployment or upgrades, impacting hospital operations. This is due to insufficient grayscale testing or incomplete rollback plans.
How to Confirm Correct Configuration
- Upload simulated medical record texts containing typical adverse reaction cases. Check if FastGPT correctly parses them and builds the knowledge base index.
- Ask questions containing medical abbreviations and professional terms. Observe if the AI accurately understands and provides relevant answers. Evaluate the effectiveness of
Similarity threshold. - Simulate a large number of concurrent requests during peak hours. Monitor system response times and resource utilization to confirm if configurations like
PARSE_FILE_TIMEOUT_SECONDScan handle the load. - Verify the process of synchronizing data from the EMR system to FastGPT. Confirm that the data update frequency aligns with business requirements.
Note: The values provided are common starting points. Measure performance against your own samples to determine optimal settings.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.