Data Characteristics for this Category
Medical insurance access documents primarily include the National Medical Insurance Catalog (covering Class A and B drugs, diagnostic and treatment items, medical services and facilities), provincial supplementary catalogs, medical insurance payment standards, negotiated drug agreements, relevant policies and regulations, clinical guidelines, pharmacoeconomic evaluation reports, and medical insurance access status of similar products domestically and internationally. Data sources are diverse, including government official websites, specialized databases, industry association publications, and academic journals. Regarding update frequency, the National Medical Insurance Catalog is typically adjusted annually. Provincial catalogs or payment standards may be dynamically adjusted based on local policies. Negotiated drug agreements have specific periodic cycles. Document structures are complex, containing both structured table data (e.g., payment scope, limited payment conditions) and extensive unstructured text (e.g., policy interpretations, clinical evidence). Fields and units involve generic drug names, dosages, specifications, medical insurance payment prices, reimbursement ratios, indications, and limiting conditions. Inconsistencies exist in data formats from different sources and unit conversion discrepancies.
Constraints Imposed by these Characteristics on "Knowledge Base Retrieval and Recall"
The complexity of medical insurance access documents imposes multiple constraints on knowledge base retrieval and recall. First, multi-source heterogeneous data makes data cleaning and standardization difficult, affecting vectorization quality and retrieval accuracy. The mixture of structured and unstructured information, in particular, requires the knowledge base to effectively integrate and process it. Second, the varying update frequencies of policies, regulations, and catalogs necessitate an efficient incremental update mechanism for the knowledge base to ensure the timeliness and accuracy of retrieval results. Outdated information retrieval due to delayed updates can impact declaration strategies. Third, the extensive text descriptions of limited payment conditions and indications contain specialized terminology and abbreviations. This requires strong semantic understanding capabilities to avoid insufficient "literal matching." Finally, medical insurance policies vary significantly across provinces. Retrieval must accurately identify regional information to avoid cross-regional policy confusion. This demands detailed design in knowledge chunking and metadata management to support multi-dimensional filtering and context awareness.
Configuration Settings
| Configuration Item | Recommended Value | Rationale for this Value |
|---|---|---|
Chunk size | 500-800 characters | Medical insurance policy texts often contain long logical chains and limiting conditions. Shorter chunks lose context, while longer ones introduce noise. |
Recall count | 8-12 entries | This covers multi-faceted policy clauses and related cases, ensuring comprehensive information and avoiding the omission of critical limiting conditions. |
Similarity threshold | 0.75-0.85 | Medical insurance clauses are precisely worded. A high threshold ensures retrieval results are highly relevant to the query intent, reducing interference from irrelevant policies. |
Rerank result count | 3-5 entries | Based on initial retrieval, a reranking model further filters the most relevant items, improving the precision of the final results. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Some medical insurance documents (e.g., detailed economic evaluation reports) can be large, requiring longer parsing times to avoid timeouts. |
Three Common Mistakes
- Retrieval results contain a large number of irrelevant expired policies or local documents. This occurs when the knowledge base does not effectively manage and filter document metadata, leading to retrieval without time range or regional attribute limitations.
- Queries for specific limited payment conditions fail to recall relevant clauses. The symptom is that returned document content literally matches keywords but fails to capture semantic-level restrictions. This is due to insufficient semantic understanding of medical insurance specialized terminology by the vector model.
- Knowledge base training progress is abnormal or stalled for a long time. This happens when PDF files containing a large number of scanned images or complex tables are uploaded. The document parser cannot effectively extract text content, leading to training failure.
How to Confirm Proper Configuration
- Select multiple typical medical insurance access query problems. Compare retrieval results to confirm the inclusion of key policy clauses, payment scopes, and limiting conditions. Evaluate their timeliness.
- Perform cross-queries for medical insurance policies across different provinces. Confirm the knowledge base accurately filters corresponding regional policy information based on metadata.
- Use query statements containing specific specialized terminology and abbreviations. Check if retrieval results accurately identify and recall relevant documents. Evaluate semantic understanding capabilities.
The values given are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.