Data Characteristics in this Domain
Medical insurance settlement data primarily originates from hospital information systems (HIS), laboratory information systems (LIS), picture archiving and communication systems (PACS), and medical insurance settlement platforms. Data updates are frequent, typically batched daily or weekly, with few real-time data interfaces. The document structure is mainly structured tabular data, including fields such as patient basic information, diagnosis, treatment, medication details, expense lists, and medical insurance payment categories. Medication details usually include generic name, brand name, dosage form, specifications, manufacturer, batch number, dosage, frequency, and administration route. The expense list specifies the unit price, quantity, and total amount for each medical service and drug. Some data may exist as unstructured text, such as handwritten doctor's notes and examination report descriptions. Field units strictly follow national medical insurance coding and clinical medical standards; for example, drug dosages are in milligrams (mg) or milliliters (ml), and frequencies are expressed as "times per day" or "times per week."
Constraints Imposed by these Characteristics on Multi-turn Conversations and Prompts
The highly structured nature of medical insurance settlement data requires precise matching of fields and values in multi-turn conversations. This demands high accuracy from the large language model (LLM) in understanding structured queries. High update frequency necessitates that the knowledge base can quickly synchronize incremental data to avoid referencing outdated information in conversations. The large number of specialized terms and codes, such as ICD-10 disease codes and ATC drug classification codes, requires prompts to guide the model in correct identification and association. Failure to do so can lead to semantic deviations or matching failures. The complex logic and multi-level categorization of expense settlements, such as reimbursement ratios, out-of-pocket ratios, and deductibles, require the model to handle complex conditional judgments and logical reasoning in multi-turn interactions to accurately answer patient or healthcare professional questions about expense composition and payment. The presence of unstructured medical record text challenges the model's entity recognition and information extraction capabilities, requiring prompts to guide the model to extract key pharmacovigilance information from text in addition to structured queries.
Configuration Settings
| Configuration Item | Suggested Value | Rationale |
|---|---|---|
maxContext | 8–12 turns | Medical insurance settlement queries often involve multiple steps of confirmation and clarification, requiring a longer conversation history to maintain context. |
Chunk size (Segment Length) | 500–800 characters | Medical insurance records are detailed; longer segments help retain complete settlement entries and medication information. |
Recall count (Recall Count) | 8–12 items | Complex medical insurance rules and medication history may involve multiple relevant records; increasing recall count improves coverage. |
Similarity threshold (Similarity Threshold) | 0.78–0.85 | Ensures recalled medical insurance data is highly relevant to the user's query, avoiding interference from irrelevant records. |
Rerank result count (Reranked Return Count) | Top 5 items | After reranking, the top few most relevant records are sufficient for the model to provide accurate answers. |
temperature | 0.3–0.5 | Medical insurance settlement questions emphasize accuracy and factuality; a low temperature helps reduce model hallucination. |
Three Common Mistakes
- The model repeatedly asks for information already provided or fails to give clear settlement basis in conversations. The
contextfield in the logs is empty or contains a large amount of duplicate information. This is becausemaxContextis set too short, leading to loss of historical conversation information and the model's inability to maintain effective context. - When a user queries about drug adverse reactions, the model only returns basic drug information and fails to link to warning events or relevant regulations. The model's answer is too broad, lacking specific risk warnings or treatment suggestions. This is because the prompt fails to effectively guide the model to retrieve specific pharmacovigilance rules and cases from the knowledge base.
- When querying medical insurance reimbursement ratios, the model gives incorrect or incomplete numbers. The numbers output by the model do not match actual reimbursement policies. This is because medical insurance policy data in the knowledge base is not updated in time, or key numbers and conditions are fragmented during data segmentation.
How to Verify Correct Configuration
- Select a series of medical insurance settlement scenarios involving multi-turn follow-up questions. Test whether the model can continuously understand the context and provide coherent answers, and compare the accuracy and completeness of the answers.
- Input queries involving specific drug adverse reactions. Verify whether the model can accurately identify drugs and link them to relevant warning information or contraindications. Check if the answer includes key risk factors.
- Test medical insurance reimbursement questions for different types of diseases and medication plans. Verify whether the model can correctly calculate reimbursement amounts, explain reimbursement policies, and compare them with actual medical insurance regulations to determine thresholds.
Note: The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.