Document Parsing and Chunking for Medical Insurance Settlement Registration and Declaration Data

Medical insurance settlement data originates from hospital HIS systems, medical insurance bureau information platforms, and pharmaceutical company

Data Characteristics

Medical insurance settlement data originates from hospital HIS systems, medical insurance bureau information platforms, and pharmaceutical company internal finance and compliance departments. Data updates typically occur monthly or quarterly, with unscheduled updates possible during medical insurance policy adjustments. Document structures vary, including standardized XML or JSON data packets, structured Excel reports, unstructured PDF policy documents, and scanned images. Key fields include generic drug names, dosages, specifications, manufacturers, medical insurance payment standards, reimbursement ratios, settlement cycles, and expense codes. Units include common currency units (Yuan), drug packaging units (boxes, bottles, sticks), dosage units (mg, g, ml), and time units (days, months, years).

Constraints on Document Parsing and Chunking

The diverse data sources for medical insurance settlement data require the document parsing module to handle both structured and unstructured data. Table structures in Excel reports need precise identification to ensure the correct association between critical figures like medical insurance payment standards and corresponding drug information. Unstructured PDF policy documents require high-precision OCR and semantic understanding to prevent policy interpretation errors due to garbled Chinese characters or recognition mistakes. The uncertainty of update frequency means chunking strategies must balance efficiency and accuracy, capable of quickly processing incremental data. The precise matching of values and units for fields like medical insurance payment standards dictates the granularity of document chunking; overly coarse chunks can lead to critical information loss, while overly fine chunks increase retrieval redundancy.

Configuration Settings

Configuration ItemRecommended ValueRationale
UPLOAD_FILE_MAX_SIZE100 MBMedical insurance policy documents and scanned images can be large, requiring support for large file uploads.
chunk_overlap50–100 charactersEnsures contextual continuity, preventing policy clauses or table data from being truncated at chunk boundaries, which affects semantic integrity.
text_splitter_nameRecursiveCharacterTextSplitterSuitable for processing multiple document types, intelligently splitting based on different delimiters, adapting to the diversity of medical insurance data.
ocr_engine_modeAccurateNumbers and text in medical insurance policy documents demand high recognition accuracy to avoid data discrepancies due to OCR errors.
min_chunk_size150 charactersEnsures each chunk contains sufficient information, avoiding the generation of too many meaningless short chunks.
table_parsing_strategyAsTextAndEmbedMedical insurance table data needs to retain both table structure information and text content for subsequent retrieval and understanding.

Common Mistakes

  • Symptom: Chinese policy clauses in PDF documents are recognized as garbled characters, or table data fields are empty after recognition. Reason: The OCR engine is not correctly configured with language packs or table recognition is not enabled, leading to unstructured data processing failure.
  • Symptom: When retrieving medical insurance payment standards, results only include numerical values without corresponding drug names or dosages. Reason: Document chunking granularity is too coarse, placing drug names and payment standards in different chunks, preventing simultaneous recall during retrieval.
  • Symptom: After uploading an Excel table, the knowledge base does not correctly display the relationships between table data, or queries for specific drug reimbursement ratios yield inaccurate results. Reason: The system parses and stores table data as ordinary text instead of structured content.

Verification Steps

  • Upload typical medical insurance policy PDF files and Excel reports. In the knowledge base management interface, check chunk content to confirm that key fields (e.g., medical insurance payment standards, drug names) are correctly extracted and contextually complete.
  • Randomly select several complex medical insurance-related queries. In the debugging interface, review the recalled chunk content to verify if it contains the core information required for the query and check the logical continuity between chunks.
  • Test scanned documents with high OCR difficulty via API or interface. Check text recognition accuracy, especially for numbers and special symbols, and compare with the original files to confirm no garbled Chinese characters.

The values provided are common starting points and should be measured against the reader's own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.