Document Parsing and Chunking for Pharmacoeconomics Quality Documents

Pharmacoeconomics documents originate from policy files and procurement documents published by national healthcare security administrations, health

Data Characteristics

Pharmacoeconomics documents originate from policy files and procurement documents published by national healthcare security administrations, health commissions, and provincial/municipal drug procurement platforms. They also include pharmacoeconomic evaluation reports submitted by pharmaceutical companies. These documents update frequently, especially during adjustments to medical insurance catalogs and the release of drug procurement plans. Document structures are complex, often containing numerous tables, charts, cited literature, and statistical data. Fields include generic drug names, dosages, specifications, medical insurance payment standards, procurement prices, drug costs, efficacy indicators (e.g., QALY, LYG), cost-effectiveness ratios (ICER), and sensitivity analysis parameters. Units vary, including RMB, USD, years, months, grams, and milligrams.

Constraints on Document Parsing and Chunking

Pharmacoeconomics documents contain dense tabular data. Traditional text chunking methods can break table integrity, separating critical values from their units. Frequent policy updates require the knowledge base to quickly identify and integrate revisions from new versions, preventing interference from outdated information. Complex document structures, especially nested tables and chart captions, challenge document parsers. These non-text elements must be correctly associated with their context. Diverse fields and mixed units require that chunks maintain semantic integrity. For example, a snippet about "a drug's annual treatment cost is 12000 yuan" must not split "12000" and "yuan" into different chunks, which would affect subsequent retrieval accuracy.

Configuration Settings

Configuration ItemRecommended ValueRationale
Chunk size (Chunk Length)800–1200 charactersEnsures complete sentences and paragraphs while preventing overly large chunks from diluting information density. Suitable for table row data.
Chunk Overlap Length (Chunk Overlap Length)150–200 charactersMaintains contextual continuity, especially for semantic connections across table or paragraph boundaries.
CHUNK_STRATEGYChunk by Title and Chunk by Table combinedPrioritizes preserving the logical structure of sections and the integrity of tabular data in pharmacoeconomics reports.
UPLOAD_FILE_MAX_SIZE500 MBAccommodates large pharmacoeconomics evaluation reports that contain numerous charts and high-resolution images.
PARSE_FILE_TIMEOUT_SECONDS600 secondsProvides sufficient parsing time for complex PDF files containing embedded fonts and charts.
PDF_OCR_ENABLEDTrueEnsures non-text content in scanned policy documents and reports can be recognized and indexed.

Common Mistakes

  • Table data is incorrectly parsed after document upload, leading to mismatched values and units in search results. This occurs because the default chunking strategy does not adequately consider table structures, splitting row or column data.
  • The knowledge base fails to update effectively after new policy files are uploaded, returning outdated information during retrieval. This happens when version management or incremental update features are not enabled, preventing new data from effectively overwriting old data.
  • Some content is unretrievable after PDF file upload, with logs showing marker-pdf parse error or connection refused. This indicates the marker-pdf service is not correctly deployed or configured, preventing access to local or remote PDF parsing services.

Verification Steps

  • Upload a pharmacoeconomics report with complex tables and charts. Review the chunk preview to ensure each row of table data (including values and units) is preserved as one or several continuous chunks.
  • Upload a policy document with clear old and new versions. Search for key revised clauses and confirm that the content from the new version is returned.
  • Randomly select 5-10 documents from the knowledge base. Perform precise searches for key terms, drug names, and pricing data within these documents. Check the accuracy and completeness of the returned results.
  • Simulate user questions about cost-effectiveness ratios or sensitivity analysis results from pharmacoeconomics evaluation reports. Verify that the answers accurately cite key data and conclusions from the source documents.

The values provided are common starting points. Measure them against your own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.