Document Parsing and Chunking for Pharmacovigilance in E-pharmacy

E-pharmacy platforms generate pharmacovigilance data from several sources. These include drug sales records, user feedback, online consultation logs

Data Characteristics

E-pharmacy platforms generate pharmacovigilance data from several sources. These include drug sales records, user feedback, online consultation logs, drug inserts, and Adverse Drug Reaction (ADR) reports. Sales records contain batch numbers, manufacturing dates, and expiration dates. User feedback and online consultations are often unstructured text, describing drug experiences and adverse symptoms. Drug inserts are structured or semi-structured documents; they cover indications, contraindications, dosage, usage, and adverse reactions. ADR reports contain detailed patient information, drug information, adverse event descriptions, interventions, and outcomes, typically in PDF or image formats. Data updates frequently, especially user feedback and sales data, which are near real-time.

Constraints Imposed by Data Characteristics on Document Parsing and Chunking

Pharmacovigilance data from e-pharmacy platforms, particularly user feedback and ADR reports, are largely unstructured. This demands advanced document parsing capabilities. Users often describe adverse reactions using colloquial language, lacking standardized medical terminology, which complicates information extraction. Drug inserts are more structured but require fine-grained chunking strategies due to their length and multi-level hierarchy. This ensures critical information like adverse reactions and contraindications are not diluted or missed. Scanned or image-based ADR reports require Optical Character Recognition (OCR) capabilities. High data update frequency necessitates an efficient incremental update mechanism for the knowledge base, avoiding frequent full re-parsing. Accurate identification of key fields, such as drug batch numbers and expiration dates, directly impacts recall precision and requires robust parsing.

Configuration Settings

Configuration ItemSuggested ValueRationale
Chunk size (Chunk Length)800–1200 charactersBalances the completeness of user feedback with the detail in ADR adverse event descriptions. Avoids excessive length, which leads to information redundancy, and insufficient length, which causes context loss.
Chunk Overlap Length (Chunk Overlap Length)100 charactersEnsures contextual coherence at chunk boundaries, especially for related information in drug inserts.
UPLOAD_FILE_MAX_SIZE100 MBAccommodates large PDF-format ADR reports or drug inserts, preventing upload failures due to excessive file size.
PARSE_FILE_TIMEOUT_SECONDS300 secondsProvides sufficient time to process complex documents containing large amounts of text or requiring OCR, reducing parsing timeouts.
Similarity threshold (Similarity Threshold)0.75Balances recall precision and coverage. Ensures effective identification of relevant adverse reaction information while filtering out noise.
Recall count (Number of Retrieved Chunks)Top 5 entries (Top 5)Limits the number of retrieved results while maintaining information richness. Improves efficiency for downstream processing and focuses on core adverse reactions.

Common Pitfalls

  • Parsing large PDF or image-format ADR reports results in "offset out of range" or "network error" messages. This typically occurs when the file size exceeds server or client transfer limits, or due to parsing timeouts.
  • Colloquial descriptions in user feedback are not correctly identified as adverse reactions, leading to missing information. This happens because the word segmentation or entity recognition model has insufficient capability to process non-standardized text, or is not optimized for medical colloquialisms.
  • Specific contraindications or precautions in drug inserts are not effectively recalled. This may be due to an overly coarse chunking strategy, where critical information is mixed with large amounts of general content, reducing its weight after vectorization.

Verification Steps

  • Upload a typical large ADR report PDF file. Check if parsing completes successfully and observe the parsing logs for errors.
  • Select multiple user feedback entries containing colloquial adverse reactions. Use knowledge base Q&A to verify accurate recall of relevant drug information and potential adverse reaction descriptions.
  • Randomly select specific contraindications or precautions from drug inserts. Verify that keyword or semantic queries can precisely locate the corresponding chunked content.
  • Continuously monitor the incremental update effectiveness of the knowledge base. Ensure newly uploaded sales data or user feedback is promptly parsed and incorporated into the knowledge base, verifying the stability of the update mechanism.

The values provided are common starting points. Measure them against your own samples for optimal results.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.