Data Characteristics
Medical insurance access product data primarily originates from policy documents, drug catalogs, diagnosis and treatment project catalogs, payment standards, and negotiation results released by national and local medical insurance bureaus. This data updates frequently. National-level adjustments typically occur annually or quarterly. Local policies may update more often, and temporary notices can be issued at any time. Document formats vary, including official announcements, formal documents (e.g., PDFs), structured tables (e.g., Excel drug catalogs), and policy interpretation texts. Data fields are complex and highly specialized, encompassing generic drug names, brand names, dosages, specifications, medical insurance payment scope, reimbursement ratios, limited payment conditions, manufacturers, and negotiated prices. Units include monetary values (Yuan), quantities (milligrams, tablets, units), and time (years, months, days).
Constraints Imposed by These Characteristics on Multi-Turn Conversations and Prompts
The high update frequency of medical insurance access data requires the knowledge base to quickly synchronize and update information, ensuring timeliness in multi-turn conversations. Document diversity, especially the mix of structured and unstructured data, challenges document processing accuracy, requiring refined handling to extract key information. Highly specialized fields and complex limiting conditions mean questions often involve multiple dimensions, such as "What are the limited conditions and reimbursement ratio for a certain drug's medical insurance payment in a specific province?" This demands that the multi-turn conversation system understands and integrates multiple contextual pieces of information, accurately linking different fields. Due to the strictness of policy details, prompt design must precisely guide the model to extract core facts, avoiding vague or inaccurate answers. Accuracy of numbers is crucial, especially when dealing with amounts and reimbursement ratios.
Configuration Guidelines
| Configuration Item | Suggested Value | Rationale |
|---|---|---|
maxContext | 6 | Medical insurance policy queries often require reviewing drug names, regions, and times from previous turns. |
Recall Count | 8–12 | Ensures coverage of key information points potentially scattered across policy documents. |
Similarity Threshold | 0.75 | Medical insurance terminology is specialized and precise, requiring a high similarity match for accuracy. |
Chunk Length | 500 characters | Policy document paragraphs are often long; this length helps retain contextual integrity. |
Rerank Count | 3 | After initial recall, selects the most relevant snippets to provide accurate answers. |
temperature | 0.3 | Medical insurance inquiries require factual answers; a lower temperature avoids generative bias. |
Common Pitfalls
- Chat history does not take effect in the workflow, leading to incoherent multi-turn conversations. This occurs when the
maxContextparameter in the workflow is set to0or historical messages are not passed correctly. - When querying the reimbursement ratio for a specific drug, results are missing or inaccurate. This happens when the knowledge base fails to accurately parse the reimbursement ratio field or limiting conditions from policy documents.
- When a user asks "What is the medical insurance payment scope for a certain drug in Shanghai?", the system cannot link to policies from a specific year. This indicates that the year was not extracted and indexed as critical metadata during document processing.
Verification Steps
- Perform multi-turn conversation tests to verify the system's ability to remember and utilize drug names and regional information mentioned in previous turns.
- Randomly select drugs from national and local medical insurance catalogs. Query their payment scope and reimbursement ratios, then cross-reference the results with original policy documents for consistency.
- Simulate complex user questions with limiting conditions, such as "What are the reimbursement regulations for a certain drug under a specific disease?", and check if the system can accurately extract and integrate relevant policy clauses.
Note: The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.