Data Characteristics
Medical insurance settlement pharmacovigilance data primarily originates from medical institution settlement documents, drug procurement details, patient medication records, and adverse event reports. This data typically exists in structured or semi-structured document formats, such as PDF settlement statements, Excel or CSV drug detail tables, and Word or RTF adverse event reports. Data update frequency is high; some settlement data may update daily or weekly, while adverse event reports generate in real-time as events occur. Document structures include settlement statements with fixed headers and multiple line items, and drug detail tables presented in tabular form. Fields include drug codes, names, specifications, dosages, costs, and settlement types. Adverse event reports may contain free-text descriptions, patient basic information, and medication history. Costs are denominated in RMB, and dosages typically involve common medical units like milligrams (mg), grams (g), and milliliters (ml).
Constraints on Document Parsing and Chunking
The high update frequency of medical insurance settlement data requires efficient and automated document parsing to handle continuous data influx. Its semi-structured nature necessitates flexible parsing strategies to accurately extract structured fields from tables and key information from free-text descriptions. Multi-line item data in settlement documents requires careful chunking granularity to ensure the completeness of individual drug or service records, preventing critical information truncation. Free-text descriptions in adverse event reports require more refined text chunking strategies to capture key entities like drug names, adverse event descriptions, and occurrence times, while avoiding context loss. Unit information (e.g., mg, ml, RMB) must be correctly identified and retained during parsing for subsequent quantitative analysis and anomaly detection.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
chunk_size | 800-1200 characters | Ensures completeness of individual line items in medical insurance settlement documents or core event descriptions in adverse event reports, while balancing recall efficiency. |
overlap_size | 100-200 characters | Maintains contextual continuity between chunks, especially when processing free-text descriptions, aiding in understanding cross-chunk information. |
parser_type | unstructured_file | Suitable for parsing multi-format documents like PDF and Word, effectively handling mixed tabular and free-text medical insurance settlement data. |
max_file_size_mb | 100 MB | Sets a reasonable upper limit for uploaded file size, considering that medical insurance settlement documents and drug detail tables may contain large amounts of data. |
metadata_extraction | enabled | Extracts metadata such as date, source, and document type, facilitating subsequent filtering and traceability. |
table_parsing_strategy | auto | Automatically identifies and parses table structures within documents, ensuring accurate extraction of structured data like medical insurance settlement details. |
Common Mistakes
- Uploading large PDF medical insurance settlement documents results in a 404 error from the document parsing node. This often occurs due to incorrect file upload path configuration after frontend packaging and deployment, preventing the server from accessing the uploaded file.
- Parsed chunk content from medical insurance settlement documents loses drug cost or dosage units. This happens when the parser fails to correctly identify and retain unit characters after numbers. Review parsing rules for unit recognition configuration.
- When processing adverse event reports, critical drug names or adverse event descriptions are truncated across different chunks. This may be related to an excessively small
chunk_sizesetting, causing individual key information to be split.
Verification Steps
- Upload a multi-page PDF medical insurance settlement statement. Check if the parsed chunks maintain the integrity of each settlement record, including all fields and units.
- Upload a Word document of an adverse event report. Verify if the parsed results accurately extract drug names, adverse event descriptions, and report dates, and that key entities are not unreasonably split.
- Randomly select parsed chunks and test them using FastGPT's query interface. Observe the relevance of recall to ensure that returned chunks effectively answer questions about specific drugs or adverse events.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.