Data Characteristics
Medical insurance settlement regulation data originates from official documents published by national and local medical insurance bureaus. These include policy notices, implementation rules, operational procedures, fee schedules (e.g., CPT/HCPCS, DRG/DIP grouping schemes), and service agreements between designated medical institutions and insurance departments. Document updates are frequent, typically following policy adjustments, with monthly or quarterly updates common and major revisions annually. Document formats are primarily PDF, Word, and HTML, containing numerous clauses, detailed rules, charts, and complex logic. Key fields include medical insurance payment scope, reimbursement ratio, deductible, cap, medical insurance catalog code, disease diagnosis code, and surgical procedure code. Numerical units involve RMB amounts, percentages, dates, and international standard codes for disease diagnosis and procedures.
Constraints on Knowledge Base Retrieval
The frequent updates to medical insurance settlement documents require the knowledge base to quickly synchronize with the latest policies. Failure to do so can lead to recalled results that do not match current regulations. The extensive use of specialized terminology and coding systems in documents demands high semantic understanding from vector models to accurately identify and associate concepts expressed in different ways. The strong logical connections between clauses mean that single-segment retrieval may be insufficient for complex questions, requiring multi-hop or chained retrieval capabilities. Documents are often lengthy and include non-textual information like tables and images, posing challenges for knowledge segmentation strategies and multimodal information extraction. Retrieval results must precisely point to the original source to allow engineers to verify policy details, necessitating accurate source document localization.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Segment Length) | 500–800 characters (characters) | Medical insurance clauses are logically dense. Shorter segments risk semantic fragmentation; longer segments introduce too much irrelevant information. |
Chunk Overlap Rate (Segment Overlap Rate) | 10%–15% | Ensures contextual continuity between adjacent segments, preventing critical information from being cut off. |
Recall count (Number of Retrieved Items) | Top 5–8 entries (top 5–8 items) | Medical insurance questions often involve multiple policy aspects; increasing retrieval count enhances coverage. |
Similarity threshold (Similarity Threshold) | Calibrate by actual measurement (calibrate by empirical testing) | Requires fine-tuning for medical insurance terminology to ensure high relevance and avoid over-generalization. |
Rerank result count (Number of Reranked Items) | Top 3 entries (top 3 items) | After processing by the reranking model, a small number of results most relevant to the user's intent are selected. |
Vector Model (Vector Model) | text-embedding-ada-002 or bge-large-zh | Selects high-performance general or Chinese-optimized models for Chinese medical insurance terminology and complex semantics. |
Common Pitfalls
- Knowledge base search tests pass, but actual question answering results in errors or inaccuracies. This may be due to an incorrectly configured
embeddingmodel in the knowledge base, leading to a mismatch between vectorized data and query vectors. - Switching the knowledge base
embeddingmodel takes a long time or fails to revert to the original model, with the UI showing "processing" or "loading failed." This usually indicates a backend task queue blockage or incorrectDBstate updates during model switching. - Retrieval results contain a large amount of irrelevant or duplicate information, or critical information is missing. This often stems from improper
Chunk size(segment length) settings, leading to incomplete semantics or information redundancy, or fromRecall count(number of retrieved items) orSimilarity threshold(similarity threshold) not being carefully tuned.
Verification Steps
- Test knowledge base question answering with typical medical insurance settlement questions. Check if the returned
referencesinclude correct and complete policy text segments. Verify thatdocument sourcepoints to the accurate file. - Use FastGPT's
knowledge base search testfeature. Input medical insurance professional terms and policy clauses. Examine thesimilarityscore distribution of theretrieved resultsand manually assess the accuracy and relevance of the retrieved content. - Regularly update a portion of medical insurance policy documents. Observe if the knowledge base
index rebuildingtask completes successfully. Immediately test relevant policy questions after updates to ensure new policies are correctly retrieved.
The values provided are common starting points. They should be measured against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.