Data Characteristics
Data for pharmacovigilance in medical insurance access primarily originates from the National Healthcare Security Administration and provincial/municipal medical insurance bureaus. This includes medical insurance drug catalogs, negotiation results, payment standard documents, drug registration approvals, package inserts, and post-market safety surveillance reports. These documents are typically PDFs or Word files, containing regulations, technical guidelines, or structured tables. Data update frequency varies; policy documents may update annually, while some drug safety information releases are irregular, based on monitoring results. Documents include fields such as generic name, trade name, ATC classification, indications, dosage and administration, adverse reactions, contraindications, precautions, medical insurance payment scope, and payment restrictions. Units include dosage (mg, g, IU), frequency (times/day, week), and cost (CNY).
Constraints Imposed by Data Characteristics on Multi-Turn Conversations and Prompts
The highly structured and policy-sensitive nature of medical insurance access data requires multi-turn conversations to accurately identify key information like drug names, restrictions, and payment scope when understanding user intent. The formal and rigorous nature of policy documents necessitates prompt design that focuses on fact extraction and rule matching, avoiding vague or speculative answers. Frequent policy updates demand strict timeliness for the knowledge base, requiring the conversation system to quickly synchronize with the latest medical insurance catalog and payment standards. Furthermore, the complexity of medical insurance payment restrictions, such as rules for specific diseases, populations, or treatment courses, requires the conversation system to perform multi-conditional logical judgments and clearly present these constraints to ensure answer accuracy.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size | 500–800 characters | Policy text paragraphs are often long; shorter segments might break complete clauses, while longer segments increase recall noise. |
Recall count | Top 8–12 entries | Medical insurance policies are highly interconnected; comprehensive judgment requires covering multiple relevant clauses. |
Similarity threshold | 0.78–0.85 | Ensures the precision of recalled content, preventing interference from irrelevant policy clauses. |
maxContext | 3500 characters | Medical insurance access questions often involve multiple follow-up questions and restrictions, requiring a longer context window. |
temperature | 0.1–0.3 | Reduces the randomness of generated answers, ensuring responses are based on knowledge base facts and align with policy rigor. |
Prompt Mode | Extractive Question Answering | Prioritizes extracting answers directly from the original text, reducing model's free generation and ensuring accurate policy interpretation. |
Common Pitfalls
- The chatbot returns inaccurate medical insurance payment information because the knowledge base data is outdated, leading the model to cite expired policy documents.
- After a user query, the conversation system responds with a long delay or "cannot answer." This may be due to
maxContextbeing set too small, truncating context information for complex queries. - API call results do not match platform test results. This is because
appIdoruserIdwere not correctly passed during the API call, preventing the loading of a specific application's knowledge base or user session history.
Verification Steps
- Create test cases based on the latest medical insurance catalog to verify if the system can accurately answer questions about drug payment scope and restrictions.
- Simulate multi-turn conversations, asking complex medical insurance policy questions involving various restrictions. Check if the system maintains contextual coherence and provides correct conclusions.
- Check logs for warning messages indicating insufficient or excessive recalled content due to improper
Chunk sizeorRecall countsettings. - Compare API call results with platform tests. Ensure that for the same input, the
answerandreferencereturned by the API match the display on the platform interface.
The values provided are common starting points and should be measured against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.