Data Characteristics
Health insurance claim registration and declaration materials primarily originate from healthcare institutions, medical insurance bureaus, and internal business systems of pharmaceutical companies. Data updates are relatively stable, typically occurring quarterly or annually in batches. This data includes policies, regulations, drug catalogs, diagnostic and treatment item catalogs, fee codes, and settlement rules. Document structures vary, encompassing PDF policy documents, Word declaration instructions, Excel fee lists and detailed tables, and structured database settlement records. Fields and units exhibit strong industry-specific characteristics, such as generic drug names, dosage forms, specifications, manufacturers, medical insurance payment standards, out-of-pocket ratios, reimbursement scopes, and settlement codes. These contain numerous codes and terminology. Standards can differ across regions and years.
Constraints from these Characteristics on "Vector Models and Indexing"
The update frequency of health insurance claim data means vector index rebuilding does not need to be frequent, but incremental updates are essential. Diverse document structures require vector models to robustly parse various document types, especially for effective extraction from table data and unstructured text. Industry-specific fields, units, codes, and terminology demand high semantic understanding from vector models. Models must accurately identify and differentiate similar medical terms and understand their specific meanings within the context of health insurance claims. Regional and annual standard differences require the index to support multi-dimensional filtering for accurate and timely retrieval results. The presence of numerous codes necessitates that vector models effectively handle these short texts during embedding to avoid information loss.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale for Recommendation |
|---|---|---|
Chunk size (Segment Length) | 500–800 characters (characters) | Balances the completeness of medical insurance policy clauses with retrieval efficiency. Avoids noise from overly long segments and loss of context from overly short segments. |
Overlap Length | 100–150 characters (characters) | Ensures semantic continuity between segments, especially in policy clauses and explanatory texts. Helps capture cross-paragraph related information. |
Vector Model (Vector Model) | bge-large-zh-v1.5 | This model performs well in Chinese semantic understanding and short text encoding, suitable for handling professional terminology and codes in health insurance claims. |
Recall count (Recall Count) | 10–15 entries (items) | Ensures coverage of initial recall results, providing sufficient candidate information for subsequent re-ranking and filtering. |
Similarity threshold (Similarity Threshold) | 0.75–0.85 | Balances recall accuracy and relevance. Avoids recalling irrelevant content while retaining enough potential matches. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds (seconds) | Addresses potentially long parsing times for large PDF or Excel files, preventing parsing failures due to timeouts. |
Common Pitfalls
- When importing many Excel files, some data rows might not be correctly parsed and vectorized. This can happen if
Chunk size(Segment Length) is too small, truncating Excel cell content, or ifPARSE_FILE_TIMEOUT_SECONDSis insufficient for large tables. - Retrieval results show numerous policy clauses irrelevant to health insurance claim rules. This indicates that
Similarity threshold(Similarity Threshold) is too low, failing to effectively filter out generalized information. - After updating the medical insurance catalog, new drug information cannot be retrieved by keywords. This indicates that the knowledge base was not incrementally updated or re-indexed, leading to inconsistencies between index data and source data.
Verification Steps
- Select representative multi-type health insurance claim documents. Manually import them and verify that all text content is correctly segmented and indexed, paying close attention to tables and complex graphic-text structures.
- For core fields like medical insurance payment standards and drug codes, use different query terms for retrieval. Observe if recall results include the expected accurate entries and check if
Recall count(Recall Count) matches the configuration. - After a policy update, perform an incremental indexing operation. Use key information from the new policy to conduct searches. Confirm that new data has been successfully incorporated into the index and can be effectively recalled.
The values provided are common starting points. Measure against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.