Data Characteristics for this Category
Medical insurance claim data primarily originates from Hospital Information Systems (HIS) and medical insurance administration business systems. Data updates are frequent, typically in daily or monthly batches, with some real-time settlement data. Document formats vary, including structured XML or JSON for settlement statements and expense details, and unstructured PDFs for patient records, hospitalization records, and diagnostic certificates. Structured data fields include patient basic information, diagnosis codes (e.g., ICD-10), surgical codes, drug and consumable codes (e.g., national medical insurance catalog codes), expense categories, payment ratios, out-of-pocket amounts, and medical insurance payment amounts. Unstructured documents often contain extensive medical terminology, abbreviations, and clinical descriptions. Fields and units may be inconsistent; for example, drug dosage units might be milligrams (mg), grams (g), or International Units (IU), while expense units are typically in Chinese Yuan.
Constraints Imposed by these Characteristics on "Model Access and Configuration"
The heterogeneous nature of medical insurance claim data requires a layered approach to model access. Structured data demands precise field mapping and type validation, such as matching medical insurance codes, which requires high accuracy in data cleaning and preprocessing. Complex medical terminology and non-standard expressions in unstructured documents necessitate stronger natural language understanding capabilities, such as named entity recognition and relationship extraction, to extract key diagnostic, treatment, and cost information from text. High update frequency requires near real-time synchronization mechanisms for the knowledge base to prevent the model from making decisions based on outdated information. Additionally, since the data contains a large amount of sensitive patient information, data anonymization and access control become mandatory constraints during model access to ensure data security and compliance. Scanned patient records in image format require image recognition capabilities to extract text content.
Determining Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Chunk Size) | 500–800 characters | Balances semantic completeness of paragraphs in medical insurance documents with model context window limitations. |
Recall count (Recall Count) | 8–12 items | Ensures sufficient relevant document chunks are covered for complex queries, avoiding omission of critical medical insurance policies or settlement rules. |
Similarity threshold (Similarity Threshold) | 0.75–0.85 | Medical insurance settlement rules demand high precision. A threshold that is too low introduces irrelevant content, while one that is too high might miss related information. |
Rerank result count (Reranked Return Count) | 3–5 items | Refines initial recall results through a reranking model, focusing on the most relevant medical insurance clauses or settlement details. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Processing large PDF scanned patient records or complex medical insurance expense details can be time-consuming, requiring ample parsing time. |
maxContext | 3500 characters | Ensures the large language model receives and understands sufficient context when handling medical insurance settlement issues, especially when multiple policies intersect. |
Three Common Mistakes
- Knowledge base query results do not provide precise medical insurance codes or reimbursement ratios. Generated content is generic. This occurs because key entities like medical insurance codes and drug names are not standardized, leading to inaccurate recall or the model's inability to correctly understand.
- Uploaded image-format patient records cannot be recognized and utilized by the model. The knowledge base preview shows empty image content. This occurs because an OCR (Optical Character Recognition) module is not configured or enabled, preventing text information from being extracted from images.
- A reranking model is configured, but no effect is observed in actual queries. Recall results are poorly ordered. This occurs because the reranking model's weights are improperly configured, or the reranking model's semantic space does not match the base recall model.
How to Confirm Proper Configuration
- Select typical medical insurance settlement questions, such as "What is the reimbursement ratio for a specific high-value consumable used for a certain disease?" Check if the model can accurately cite relevant medical insurance policy terms and specific figures.
- Upload medical insurance settlement documents containing images and tables. Use the knowledge base preview function to check if text from images and table data are correctly extracted and segmented.
- Simulate medical insurance policy updates by updating some knowledge base content. Then, query related questions to verify if the model can promptly reflect the latest policy changes.
- For common ambiguous queries in medical insurance settlement, such as "Is a certain drug reimbursable?", check if the recall results include various possible reimbursement scenarios and restrictions, and observe if the relevance improves after reranking.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.