Data Characteristics
Pharmacovigilance data in nursing management originates from electronic health records (EHRs), nursing notes, medication orders, adverse event reports, and periodic health assessments. This data updates frequently. Some medication records and vital sign data may update hourly or even minute-by-minute. Adverse event reports are entered immediately after an event occurs. EHRs typically use a hybrid of structured and semi-structured formats. They include explicit fields such as patient ID, medication name, dosage, administration route, and administration time. They also contain extensive free-text descriptions, such as nursing observation records and patient complaints. Adverse event reports have fixed classification and severity assessment fields, along with detailed event descriptions. Fields and units involve milligrams (mg), milliliters (ml), international units (IU) for dosage. Time units are dates and timestamps down to the minute. Vital sign data includes blood pressure (mmHg) and heart rate (beats/min).
Constraints on Document Parsing and Chunking
The high update frequency of nursing management data requires a document parsing system to rapidly process new and changed documents. This prevents data lag from affecting the real-time nature of pharmacovigilance. The hybrid structured and semi-structured document format requires balancing accurate field extraction with deep semantic understanding of free text. Free text in nursing observation records often describes patient symptoms and vital sign changes. These are crucial clues for identifying adverse reactions. Detailed chunking is necessary to preserve contextual relevance. Precise field and unit information is vital for dosage verification and adverse reaction judgment. Parsing must avoid unit confusion or misinterpretation of numerical values. Additionally, fixed classification fields in adverse event reports must have their label information indexed with related descriptive text after chunking, facilitating subsequent retrieval.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Chunk Length) | 500–800 characters | Observation descriptions in nursing records and adverse event details require a moderate length to retain complete semantic context. Too short may fragment critical information; too long introduces noise. |
Chunk Overlap Length (Chunk Overlap Length) | 100 characters | Ensures sufficient overlap between adjacent chunks. This addresses cases where critical information might appear at chunk boundaries, maintaining semantic coherence. |
maxContext | 4000 tokens | Pharmacovigilance analysis often requires correlating multiple medication records, nursing observations, and adverse event reports. A larger context window helps AI models make comprehensive judgments. |
UPLOAD_FILE_MAX_SIZE | 50 MB | Electronic health records and reports may contain charts and large amounts of text. This value accommodates the size of most medical documents, preventing upload failures. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Complex semi-structured document parsing can take a long time. Extending the timeout ensures large files can be fully processed. |
Recall count (Retrieval Count) | 8–12 items | Pharmacovigilance involves many factors. Appropriately increasing the retrieval count improves the coverage of relevant information, reducing the risk of missed reports. |
Common Pitfalls
- The system reports a file parsing failure with status code
413 Request Entity Too Large. This usually means the uploaded document size exceeds theUPLOAD_FILE_MAX_SIZEparameter limit. - The AI model fails to accurately link to specific patient medication dosages or administration times in its responses. This may occur if dosage units or time formats were not correctly recognized during document parsing, leading to missing or incorrectly parsed field values.
- When retrieving adverse reaction information, semantically related long descriptions are sometimes unreasonably split. This results in retrieved chunks lacking complete context. This typically happens when
Chunk size(Chunk Length) is set too small, orChunk Overlap Length(Chunk Overlap Length) is insufficient.
Verification Steps
- Upload representative EHRs, nursing notes, and adverse event reports. Check if the parsed chunks are semantically complete. Verify that critical information (e.g., medication name, dosage, time, symptom description) is correctly extracted and retained in independent or associated chunks.
- Use the API to retrieve the chunked index content of a document. Cross-reference the parsed results of structured fields (e.g., medication names, timestamps) with the original document. Pay close attention to unit matching.
- Simulate pharmacovigilance scenarios with test questions. Ask about adverse reaction information for specific patients. Evaluate if the AI model can accurately retrieve and integrate relevant information from different chunks.
- Continuously monitor system logs for file parsing success rates and processing times. Ensure stable processing even during peak data update periods.
The values provided are common starting points. Measure against your own data samples for optimal configuration.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.