Document Parsing and Chunking for Rare Disease Pharmacovigilance

Rare disease pharmacovigilance data originates from clinical trial reports, real-world evidence (RWE) data, individual case safety reports (ICSRs)

Data Characteristics

Rare disease pharmacovigilance data originates from clinical trial reports, real-world evidence (RWE) data, individual case safety reports (ICSRs), medical literature, drug labels, and regulatory safety updates. Document update frequencies vary. Clinical trial reports and drug labels update slowly. ICSRs and medical literature can emerge in real time. Document structures are diverse, ranging from structured electronic report forms to unstructured free-text reports. Common fields include drug name, dosage, administration route, adverse event terms (e.g., MedDRA codes), event date, and outcome. Units involve time (days, weeks, months) and dosage (mg, g, IU), with multiple possible expressions. Global data involves multilingual text and varying report formats across countries and regions.

Constraints on Document Parsing and Chunking

The diverse data sources for rare diseases require parsers to handle multiple file formats, including PDF, DOCX, XML, and plain text. Uncertain update frequencies, especially real-time ICSRs, demand high parsing efficiency and incremental update capabilities. Diverse document structures mean fixed-template parsing methods are insufficient. Flexible text extraction and information retrieval capabilities are necessary. Multilingual text and medical terminology require parsing models with language understanding and medical vocabulary recognition to ensure accurate extraction of key information. Rare disease reports may contain extensive non-standardized descriptions, requiring refined chunking strategies to avoid splitting or omitting critical information.

Configuration Settings

Configuration ItemRecommended ValueRationale
UPLOAD_FILE_MAX_SIZE500 MBAccommodates large clinical trial reports and documents with numerous attachments.
PARSE_FILE_TIMEOUT_SECONDS600 secondsEnsures complex PDFs and multi-page documents have sufficient time to parse, preventing timeouts.
Chunk size800–1200 charactersBalances context completeness with retrieval efficiency, adapting to long sentences and complex medical descriptions in rare disease reports.
Chunk Overlap Length100 charactersEnsures contextual continuity at chunk boundaries, reducing the risk of critical information being cut off.
maxContext32000 tokenAdapts to large language model input windows, especially when processing complex medical literature.
enable_ocrtrueEnsures text in scanned documents or images is recognized and parsed.

Common Pitfalls

  • A split error from the parser often indicates special characters or formatting in the document content, preventing correct text chunking.
  • An API call for file parsing succeeds but returns no result. This may be due to a parsing timeout or an unhandled error in the backend service.
  • Incomplete retrieval from some documents (e.g., Feishu documents, embedded images in DOC files) after parsing. This occurs because the default parser has limited capability in handling specific platforms or embedded objects, failing to extract all text information.

Validation Steps

  • Select rare disease pharmacovigilance documents from various sources and formats. Upload them and verify that the parsed text content is complete and free of garbled characters.
  • For documents with multi-level directories, check if all levels of content, especially deeply nested information, are retrievable after parsing.
  • Test with documents containing images or scanned text to confirm that the enable_ocr function correctly identifies and extracts text information from images.

The values provided are common starting points and should be measured against specific samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.