Document Parsing and Chunking for Pharmacoeconomic Pharmacovigilance

Pharmacoeconomic data often originates from clinical trial reports, real-world evidence (RWE) studies, medical cost analysis reports, pharmaceutical

Data Characteristics

Pharmacoeconomic data often originates from clinical trial reports, real-world evidence (RWE) studies, medical cost analysis reports, pharmaceutical company market access documents, and government drug procurement and reimbursement policy documents. These documents are primarily in PDF format. They contain substantial structured and semi-structured data, including disease burden data, treatment pathways, drug pricing, reimbursement standards, cost-effectiveness analysis models, and sensitivity analysis results. Data update frequency is relatively stable, typically aligning with new drug approvals, indication expansions, or adjustments to medical insurance catalogs. Common fields in these documents include QALYs (Quality-Adjusted Life Years), ICER (Incremental Cost-Effectiveness Ratio), drug generic names, brand names, dosages, administration routes, treatment durations, patient population characteristics, and various monetary units (e.g., USD, EUR, CNY) and time units (e.g., year, month, week).

Constraints on Document Parsing and Chunking

The characteristics of pharmacoeconomic documents impose specific requirements on document parsing and chunking. First, tables and charts are core information carriers, such as cost-effectiveness matrices and sensitivity analysis results. This means simple text segmentation is insufficient to capture complete semantics; structured extraction of table and image content is necessary. Second, documents are typically long, often hundreds or even thousands of pages, with high information density. Parsers must efficiently handle large files to avoid memory overflows or timeouts. Third, numerous specialized terms and abbreviations exist, such as QALY and ICER. Chunking must preserve the integrity of these key metrics and their context to prevent semantic loss due to truncation. Fourth, monetary and time units may vary across reports, affecting subsequent numerical comparisons and calculations. Chunking must retain this unit information. Finally, documents often contain multi-level headings and chapter structures. Parent-child chunking is crucial for maintaining logical document hierarchy, which significantly improves recall accuracy.

Configuration Settings

Configuration ItemSuggested ValueRationale
Chunk size (Chunk Size)800–1200 characters (characters)Balances semantic completeness with recall efficiency, preventing noise from overly long chunks and context loss from overly short chunks.
Chunk overlap (Chunk Overlap)100–200 characters (characters)Ensures semantic continuity between adjacent chunks, preventing critical information from being truncated.
Parent-Child ChunkingEnablePharmacoeconomic documents have complex structures with multi-level headings. Enabling this preserves the document hierarchy.
Table RecognitionEnableDocuments contain extensive tabular data from cost-effectiveness and sensitivity analyses. Enabling this allows structured extraction.
Image OCREnableExtracts key text information from charts and graphs, such as axis labels, data point descriptions, and legends.
PARSE_FILE_TIMEOUT_SECONDS600 seconds (seconds)Large PDF documents take longer to parse. Increasing the timeout ensures processing completion.

Common Mistakes

  • Parsing large PDF documents results in timeout or out-of-memory errors. This occurs because the PARSE_FILE_TIMEOUT_SECONDS configuration is too low or system resources are insufficient to process thousands of pages.
  • Data in tables or charts is not extracted correctly, leading to incomplete retrieval results. This happens when Table Recognition or Image OCR is not enabled, or the parser version does not support these features.
  • Context for key terms is missing in retrieval results, for example, QALY values and corresponding cost data are in different chunks. This is due to a Chunk size (Chunk Size) that is too short, or Parent-Child Chunking is not enabled, leading to a loss of logical structure.

Configuration Verification

  • Select several typical and representative pharmacoeconomic documents (e.g., cost-effectiveness analysis reports). Upload them and observe if parsing completes successfully.
  • Examine the parsed chunk content in the knowledge base. Focus on verifying whether tabular data, chart descriptions, and key metrics are fully extracted.
  • Perform keyword searches on the parsed chunks. Verify if relevant information is accurately recalled and evaluate the completeness of the retrieved context.

The values given are common starting points and should be measured against the reader's own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.