Data Characteristics
E-pharmacy platforms generate pharmacovigilance data from several sources. These include drug sales records, user feedback, online consultation logs, drug inserts, and Adverse Drug Reaction (ADR) reports. Sales records contain batch numbers, manufacturing dates, and expiration dates. User feedback and online consultations are often unstructured text, describing drug experiences and adverse symptoms. Drug inserts are structured or semi-structured documents; they cover indications, contraindications, dosage, usage, and adverse reactions. ADR reports contain detailed patient information, drug information, adverse event descriptions, interventions, and outcomes, typically in PDF or image formats. Data updates frequently, especially user feedback and sales data, which are near real-time.
Constraints Imposed by Data Characteristics on Document Parsing and Chunking
Pharmacovigilance data from e-pharmacy platforms, particularly user feedback and ADR reports, are largely unstructured. This demands advanced document parsing capabilities. Users often describe adverse reactions using colloquial language, lacking standardized medical terminology, which complicates information extraction. Drug inserts are more structured but require fine-grained chunking strategies due to their length and multi-level hierarchy. This ensures critical information like adverse reactions and contraindications are not diluted or missed. Scanned or image-based ADR reports require Optical Character Recognition (OCR) capabilities. High data update frequency necessitates an efficient incremental update mechanism for the knowledge base, avoiding frequent full re-parsing. Accurate identification of key fields, such as drug batch numbers and expiration dates, directly impacts recall precision and requires robust parsing.
Configuration Settings
| Configuration Item | Suggested Value | Rationale |
|---|---|---|
Chunk size (Chunk Length) | 800–1200 characters | Balances the completeness of user feedback with the detail in ADR adverse event descriptions. Avoids excessive length, which leads to information redundancy, and insufficient length, which causes context loss. |
Chunk Overlap Length (Chunk Overlap Length) | 100 characters | Ensures contextual coherence at chunk boundaries, especially for related information in drug inserts. |
UPLOAD_FILE_MAX_SIZE | 100 MB | Accommodates large PDF-format ADR reports or drug inserts, preventing upload failures due to excessive file size. |
PARSE_FILE_TIMEOUT_SECONDS | 300 seconds | Provides sufficient time to process complex documents containing large amounts of text or requiring OCR, reducing parsing timeouts. |
Similarity threshold (Similarity Threshold) | 0.75 | Balances recall precision and coverage. Ensures effective identification of relevant adverse reaction information while filtering out noise. |
Recall count (Number of Retrieved Chunks) | Top 5 entries (Top 5) | Limits the number of retrieved results while maintaining information richness. Improves efficiency for downstream processing and focuses on core adverse reactions. |
Common Pitfalls
- Parsing large PDF or image-format ADR reports results in "offset out of range" or "network error" messages. This typically occurs when the file size exceeds server or client transfer limits, or due to parsing timeouts.
- Colloquial descriptions in user feedback are not correctly identified as adverse reactions, leading to missing information. This happens because the word segmentation or entity recognition model has insufficient capability to process non-standardized text, or is not optimized for medical colloquialisms.
- Specific contraindications or precautions in drug inserts are not effectively recalled. This may be due to an overly coarse chunking strategy, where critical information is mixed with large amounts of general content, reducing its weight after vectorization.
Verification Steps
- Upload a typical large ADR report PDF file. Check if parsing completes successfully and observe the parsing logs for errors.
- Select multiple user feedback entries containing colloquial adverse reactions. Use knowledge base Q&A to verify accurate recall of relevant drug information and potential adverse reaction descriptions.
- Randomly select specific contraindications or precautions from drug inserts. Verify that keyword or semantic queries can precisely locate the corresponding chunked content.
- Continuously monitor the incremental update effectiveness of the knowledge base. Ensure newly uploaded sales data or user feedback is promptly parsed and incorporated into the knowledge base, verifying the stability of the update mechanism.
The values provided are common starting points. Measure them against your own samples for optimal results.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.