Data Characteristics
Health insurance access policy data primarily originates from official policy documents, interpretation announcements, drug catalogs, and treatment catalogs published by national and provincial medical insurance administrations. This data updates frequently, typically with national or local policy adjustments. Examples include annual adjustments to the medical insurance catalog and irregular temporary policy supplements. Documents are often formal PDF files containing legal clauses, technical specifications, and appended tables. Drug catalogs include fields such as generic drug name, dosage form, specification, medical insurance payment standard, and restricted payment scope. Units involved include dosage units (e.g., mg, g), quantity units (e.g., tablets, vials), and monetary units (e.g., CNY). Treatment items involve project codes, names, prices, and payment categories.
Constraints Imposed by These Characteristics on "Reference and Traceability"
The high timeliness of health insurance access policy data requires the knowledge base to synchronize updates quickly. This ensures that referenced policy bases are the latest versions. PDF-formatted legal documents often contain complex tables and multi-level headings. This requires document parsers to be robust, accurately extract text content, and preserve structural information to prevent context breaks during citation. The rigor of policy clauses dictates that reference sources must be precise down to the original document's chapter, clause, or even specific table row. This meets the accuracy requirements for traceability. Furthermore, correct identification and display of units for numerical data, such as medical insurance payment standards, directly impact the usability of question-answering results. For example, payment standard should clearly indicate CNY.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Segment Length) | 500–800 characters (characters) | Ensures the completeness of policy clauses, preventing critical information from being split, while balancing retrieval efficiency. |
Recall count (Recall Count) | Top 5–8 entries (top 5–8 items) | Medical insurance policies are highly interconnected. Recalling more items helps cover relevant regulations and improves answer comprehensiveness. |
Similarity threshold (Similarity Threshold) | 0.78–0.85 | Policy texts have precise semantics. Increasing the threshold reduces the occurrence of irrelevant or ambiguous references. |
Rerank result count (Rerank Return Count) | Top 3 entries (top 3 items) | After reranking, the few most relevant items usually provide core information, preventing information overload. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds (seconds) | Health insurance policy documents are often large, and parsing can be time-consuming. Sufficient time must be allocated. |
maxContext | 3000 Tokens | Ensures the model has enough context to understand and reference policy clauses when generating answers. |
Common Pitfalls
- AI dialogue output displays garbled characters like
[1]for citations, which then revert to quotation marks. This typically indicates a front-end rendering or character encoding issue. Check page encoding settings or how front-end components handle special characters. - AI provides citations for questions not present in the knowledge base. This suggests the
Similarity threshold(Similarity Threshold) is set too low. The model returns low-similarity document segments even when no highly matching content is found. - The dialogue request interface does not return
citereference IDs. This may occur if the knowledge base module in the workflow is not configured to output citation information, or if the interface's returned data structure does not include this field.
Verification Steps
- Test with specific health insurance access policy questions. Verify that the cited sources in the answers point to the correct policy documents and specific clauses.
- After uploading new health insurance policy documents, perform a retrieval and question-answering session. Confirm that the new document content is correctly indexed and cited. Verify that the
citefield contains valid file ID and page number information. - Simulate asking medical insurance questions not present in the knowledge base. Observe if the AI refuses to answer or provides a "no relevant information found" prompt without citations. This verifies the effectiveness of the
Similarity threshold(Similarity Threshold).
The values provided are common starting points and should be measured against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.