Data Characteristics
Core data for medical insurance settlement products originates from official documents. These include policy regulations, payment standards, drug catalogs, and treatment catalogs published by national and local medical insurance bureaus. Data updates frequently, typically quarterly or annually, with ad-hoc updates for significant policy changes. Documents come in various formats: PDF policy texts, Excel catalog lists, and structured database entries for codes and pricing. Key fields include disease diagnosis codes (ICD-10), surgical procedure codes, generic drug names, dosage forms, specifications, medical insurance payment categories, payment ratios, price limits, and medical service item codes and pricing. Data units include monetary amounts (Yuan), percentages (%), and quantities (boxes, tablets, times).
Constraints on Knowledge Base Retrieval
High update frequency for medical insurance settlement data requires the knowledge base to support efficient incremental updates and version management. This ensures timely and accurate retrieval results. Policy documents in PDF format and tabular data present challenges for document parsing and structured extraction. Accurate identification and extraction of key information is necessary. Diverse coding systems (e.g., ICD-10) and complex payment rules necessitate support for multi-field, multi-condition combined queries. Regional differences in medical insurance policies require the knowledge base to perform precise recall based on geographic context, preventing cross-regional information confusion. Numerical data, such as payment ratios and price limits, requires range queries and numerical comparisons during retrieval to meet business needs.
Configuration Settings
| Configuration Item | Suggested Value | Rationale |
|---|---|---|
Chunk size (Segment Length) | 500–800 characters (characters) | Medical insurance policy clauses are often long. An appropriate segment length maintains context completeness and prevents key information from being split. |
Chunk Overlap Length (Segment Overlap Length) | 100–150 characters (characters) | Ensures sufficient overlap between adjacent segments. This addresses cases where keywords span segments, improving recall rate. |
Recall count (Recall Count) | Top 5 entries (top 5) | Considering the complexity of medical insurance policies, increasing the recall count improves coverage and reduces omissions. |
Similarity threshold (Similarity Threshold) | Calibrate by actual measurement (calibrate by actual measurement) | This value requires iterative testing against actual corpus and business requirements. This ensures results are relevant without being overly broad. |
Rerank result count (Rerank Return Count) | 3 entries (3 items) | Given the specialized nature of medical insurance inquiries, reranking selects the most relevant few items. This helps users quickly obtain core information. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds (seconds) | Processing large PDF policy files requires sufficient file parsing timeout. This prevents parsing failures due to excessively large files. |
Common Pitfalls
- Retrieval results contain outdated policies or invalid payment standards. This occurs when the knowledge base does not synchronize the latest medical insurance policy documents promptly, or when version management is inadequate.
- User queries for specific disease reimbursement ratios return empty or irrelevant results. This can happen if medical insurance codes (e.g.,
ICD-10) are not correctly extracted or indexed, leading to a mismatch. - The progress bar stalls for an extended period or an error
Document parsing failedappears after document import. This usually indicates that the imported PDF file is overly complex, contains many images, or is encrypted, causing the file parser to time out or fail to recognize content.
Verification Steps
- Select recently updated medical insurance policy documents. After importing them into the knowledge base, check if key fields (e.g.,
payment ratio,price limit) are accurately extracted and retrievable. - Simulate user queries for different regions and types of medical insurance. Verify that the knowledge base recalls policy entries corresponding to the correct region and category.
- Test complex queries containing
ICD-10codes or generic drug names. Confirm that retrieval results include policy regulations strongly related to the code or drug name. - Perform import tests with medical insurance files of varying sizes and formats (PDF, Excel). Ensure file parsing and segmentation processes are error-free and knowledge base content is complete.
The values provided are common starting points. Measure them against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.