Data Characteristics
Medical insurance settlement regulation data originates from policy documents, implementation rules, operational guidelines, and related interpretations published by national and local medical insurance bureaus. These documents are typically in PDF, Word, or HTML format. Update frequencies vary; national policies may adjust annually, while local rules might update quarterly or semi-annually. Document structures often include chapter titles, clause numbers, specific content descriptions, explanatory notes, and attachments. Fields and units involve diagnosis codes (e.g., ICD-10), generic drug names (e.g., ATC classification), service facility fees, deductible amounts (unit: CNY), reimbursement ratios (unitless, typically percentage), payment limits (unit: CNY), and settlement periods (unit: days). This data exhibits strong hierarchical relationships and cross-referencing.
Constraints on Multi-Turn Conversations and Prompts
The hierarchical and cross-referencing nature of medical insurance settlement data requires precise understanding of user intent in multi-turn conversations. The system must identify entities in context (e.g., specific diseases, drugs, or medical services) and trace them back to original policy clauses. Policy update frequency demands efficient incremental update and version management capabilities for the knowledge base to ensure timely conversation content. The extensive use of specialized terminology and codes in documents places high demands on prompt semantic understanding and entity recognition capabilities. For example, when a user queries "chronic disease reimbursement," the system must identify "chronic disease" and link it to specific disease catalogs, corresponding reimbursement ratios, and deductibles. Accurate extraction and calculation of numerical fields like amounts and ratios also require prompts to guide the model toward precise numerical reasoning, avoiding vague answers.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
maxContext | 8 | Ensures coverage of multiple related policy clauses and details in multi-turn conversations while controlling token consumption. |
recall_top_k | 5 | Balances the complexity of medical insurance policies and the breadth of user queries, increasing the likelihood of recalling relevant policy segments. |
similarity_threshold | 0.78 | Medical insurance policy texts have rigorous semantics, requiring a high similarity match to ensure the accuracy of recalled content. |
chunk_size | 800–1000 characters | Policy document paragraphs are often long and contain multiple regulations. This length helps maintain semantic completeness. |
prompt_template | custom | Must include clear guidance on medical insurance policies, diagnosis codes, and reimbursement ratios, such as "Based on medical insurance policy documents and the diagnosis item provided by the user, please state the reimbursement ratio and payment limit." |
temperature | 0.3 | Medical insurance policy Q&A requires precise and objective results. A lower temperature reduces the model's generation of divergent or creative answers. |
Common Pitfalls
- Symptom: A user asks about the "reimbursement ratio for a specific drug," and the system returns the drug's reimbursement ratios for various diseases. Reason: The prompt failed to explicitly guide the model to filter based on the disease context provided by the user, or the knowledge base associated too many undifferentiated policy clauses with the drug.
- Symptom: A user queries the medical insurance reimbursement limit, and the system responds, "No relevant information found." Reason: The knowledge base failed to effectively ingest or index payment limit data presented in tabular format within policy documents, preventing the model from extracting specific values.
- Symptom: During a conversation, the user repeatedly mentions "deductible," but the system fails to maintain a consistent understanding of its referent, repeatedly asking the user for clarification. Reason: The prompt failed to effectively utilize
maxContextto maintain contextual coherence, or entity recognition and disambiguation for core concepts like "deductible" were insufficient.
Validation Steps
- Test whether multi-turn conversations can accurately identify user intent and provide consistent answers for medical insurance settlement queries of varying complexity, especially in scenarios involving amount calculations and policy clause citations.
- Randomly select recently published national and local medical insurance policy documents, import their content into the knowledge base, and test whether the system can accurately answer questions regarding key clauses. Verify the effectiveness of
recall_top_kandsimilarity_threshold. - Simulate a user repeatedly mentioning or modifying key entities (e.g., "drug name," "diagnosis code") within a conversation. Observe whether the system can correctly track and update its understanding within the
maxContextrange, avoiding context loss.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.