Data Characteristics
Medical insurance settlement data comes from various sources. These include policy documents, payment standards, drug and consumable catalogs, and treatment catalogs published by national and local medical insurance bureaus. This data updates frequently. National policies typically update annually, while local policies may adjust quarterly or irregularly. Document structures are complex, often in PDF, Word, or Excel formats. They contain numerous nested tables, legal citations, and technical specifications. Field names are inconsistent. For example, "medical service item code" may have different names across provinces. Units can also vary, such as "times," "courses," or "person-days." Some data involves critical values like medical fund payment ratios and individual co-payment ratios, often presented as percentages or fixed amounts.
Constraints Imposed by These Characteristics on Knowledge Base Retrieval and Recall
Frequent updates to medical insurance settlement data require the knowledge base to have an efficient incremental update mechanism. This ensures timely retrieval results. Complex document structures mean traditional text chunking methods may struggle to capture complete semantic information, especially regarding table data relationships. Inconsistent field names and units increase retrieval difficulty, demanding smarter entity recognition and matching capabilities. The precision required for legal provisions and technical specifications sets a higher standard for recall accuracy. This prevents misjudgments due to lost context. Furthermore, the presence of medical fund payment ratios means retrieval must consider numerical range queries and unit conversions to meet user needs for specific reimbursement conditions.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
chunk_size | 800–1200 characters | Balances policy clause completeness and retrieval efficiency |
chunk_overlap | 100 characters | Ensures context continuity and reduces semantic fragmentation |
retrieve_top_k | top 8 | Balances retrieval scope with subsequent re-ranking processing capability |
similarity_threshold | 0.78–0.85 | Filters low-relevance results, improving recall quality |
rerank_top_n | top 3 | Focuses on the most relevant key information, reducing user reading burden |
reranker_model | bge-reranker | Offers good comprehension and ranking ability for long texts and complex semantics |
Common Misconfigurations
- Symptom: Retrieval results include a large number of outdated policies or repealed clauses. Reason: The knowledge base did not update promptly, or a version management mechanism was missing, leading to ineffective old data cleanup.
- Symptom: When a user queries "reimbursement ratio for a certain item," the results fail to accurately state the specific percentage or return multiple irrelevant values. Reason: The knowledge base chunking strategy did not adequately consider the completeness of table and numerical data, separating key values from their descriptions.
- Symptom: After configuring
reranker_model, sorting remains chaotic, with irrelevant results ranked higher. Reason: The selected model's matching ability for medical insurance settlement domain-specific terminology and complex sentence structures is insufficient, or the model's weights were not optimized for this domain.
Validation
- Select a batch of typical and challenging medical insurance settlement questions. Observe whether the knowledge base's retrieved items contain the correct answers and check their ranking in the recall list.
- For questions involving numerical queries (e.g., reimbursement ratios, payment limits), verify the accuracy of numerical values in the recall results against the original documents.
- Evaluate the knowledge base's performance when processing different document types (e.g., policy documents, catalog lists, technical specifications). Ensure effective retrieval for all data structures.
- Regularly track medical insurance policy updates. After updates, verify the knowledge base's retrieval and recall capabilities for new policies and ensure old policies are correctly identified or removed.
The values provided are common starting points. Measure them against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.