Data Characteristics
Pharmacovigilance data in health management comes from individual health records, wearable device data, medical imaging reports, laboratory test results, and patient-reported medication histories. Data updates are frequent; some physiological indicators update every minute. Adverse drug reaction reports are typically entered within hours or days of an event. Document structures are diverse, including structured electronic medical record fields and extensive unstructured text like doctor's notes, patient feedback, and medication diaries. Fields and units are complex. For example, blood pressure is in mmHg, blood glucose in mmol/L or mg/dL, and drug dosages in mg, g, or IU. Conversions between different standards are often required.
Constraints on Knowledge Base Retrieval and Recall
Rapid data updates require real-time synchronization capabilities for the knowledge base. Traditional periodic batch processing can lead to outdated information. Diverse document structures mean a single text segmentation strategy is insufficient; it requires combining structured information extraction with unstructured text semantic understanding. Accurate field and unit handling is central to pharmacovigilance. Incorrect or missing unit information can lead to misinterpretations in recall results; for instance, dosage unit confusion directly impacts risk assessment. Health management data involves sensitive personal information, demanding strict anonymization and access control for recall results to protect privacy. Retrieving long-tail, low-frequency adverse event data requires the knowledge base to have strong semantic matching capabilities to precisely identify rare patterns from large datasets.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Segment Length) | 500–800 characters | Balances contextual completeness and retrieval efficiency. Avoids noise from overly long paragraphs while ensuring key information is not fragmented. |
Recall count (Number of Retrieved Items) | 8–12 items | Balances potential relevance and processing load. Ensures coverage of enough candidate information without overwhelming subsequent re-ranking. |
Similarity threshold (Similarity Threshold) | 0.75–0.85 | Addresses the precision requirements of medical terminology, improving recall accuracy and reducing irrelevant results. |
Rerank result count (Number of Re-ranked Items) | 3–5 items | Focuses on core relevant information, reducing redundant content presented to the user and improving user experience. |
MAX_FILE_SIZE_MB | 200 MB | Accommodates the import of large files like health records and imaging reports, preventing upload failures due to excessive file size. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Handles the time-consuming parsing of complex structured documents and long texts, preventing parsing timeouts. |
Common Mistakes
- When uploading large health report files, the system returns a
413 Request Entity Too Largeerror. This occurs because theMAX_FILE_SIZE_MBparameter is set too low to accommodate the file size. - Retrieval results contain many irrelevant medication records, but critical adverse reaction information is missing. This happens when the segmentation strategy is too coarse, failing to effectively identify and separate core event descriptions.
- After importing WeChat official account articles into the knowledge base, some content is not retrievable, resulting in no query results. This is because HTML tags were not effectively cleaned during import, leading to incorrect parsing or truncation of text content.
Verification Steps
- Import real health records and pharmacovigilance reports in various formats (PDF, TXT, JSON). Check if the knowledge base document count and content are complete and accurate.
- For known adverse events, use different keywords to perform retrieval. Verify the accuracy and ranking of recall results, ensuring important information appears prominently.
- Simulate high-concurrency query scenarios. Monitor system response time and resource utilization to confirm that parameters like
PARSE_FILE_TIMEOUT_SECONDSsupport actual business loads.
The values provided are common starting points and should be measured against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.