Data Characteristics
Medical insurance claims data originates from hospital information systems (HIS), including doctor's orders, patient admission records, expense lists, and detailed settlement records from medical insurance bureaus. This data is primarily structured and semi-structured, such as XML, JSON, or CSV expense lists, alongside unstructured medical record texts. Update frequency is typically daily or weekly, depending on the hospital's data synchronization mechanism with the medical insurance bureau. The document structure is complex, containing patient demographics, diagnoses, treatment plans, drug lists, consumable details, service item codes, and corresponding costs. Common fields include patient_id, diagnosis_code (ICD-10), drug_code (national medical insurance code), service_code, amount (unit: CNY), and unit (e.g., box, time, tablet).
Constraints on Knowledge Base Retrieval and Recall
The multi-source and complex nature of medical insurance claims data imposes requirements on knowledge base preprocessing and retrieval strategies. Data update frequency dictates the design of the knowledge base synchronization mechanism, which must support incremental updates to ensure timeliness. The presence of structured fields, like medical insurance codes, requires the retrieval system to support exact matching and metadata-based filtering. Unstructured medical record texts necessitate high-quality text segmentation and embedding generation to capture semantic information within clinical descriptions. Numerical fields such as amount and unit may involve range queries or numerical comparisons during retrieval, requiring embedding models to effectively encode numerical information or use metadata filtering. Large data volumes and strong field interdependencies challenge recall efficiency and accuracy, demanding fine-grained control over recall granularity.
Configuration Guidelines
| Configuration Item | Suggested Value | Rationale |
|---|---|---|
Chunk size (Segment Length) | 500–800 characters | Accommodates varying text lengths like doctor's orders and medical record summaries, ensuring single-segment information completeness and preventing critical information truncation. |
Overlap Length | 50 characters | Ensures contextual continuity and prevents semantic fragmentation due to segmentation, especially for descriptions of diagnostic and treatment processes. |
Recall count (Number of Retrieved Items) | Top 8 | Considering that medical insurance claims involve multiple expense details and diagnostic information, increasing the number of retrieved items appropriately covers potential relevant entries. |
Similarity threshold (Similarity Threshold) | Calibrated by actual measurement | Iteratively optimize using metrics like F1-score based on actual query performance, balancing recall and precision. |
embedding_model | text-embedding-ada-002 | Balances encoding efficiency and semantic representation capabilities, showing good generalization for Chinese medical texts. |
maxContext | 4096 tokens | Ensures the large language model receives a sufficiently long context to accommodate multiple retrieved medical insurance claim records and associated metadata. |
Common Pitfalls
- Missing partial expense details or diagnostic information in query results. This occurs due to an improper knowledge base segmentation strategy, leading to critical information being truncated or scattered across different segments, preventing effective recall.
- Slow retrieval speed, with response times exceeding
10 seconds. This can be due to insufficient knowledge base index optimization or setting an excessively largetop_kfor retrieved items, increasing the burden on vector retrieval and post-processing. - The system returns irrelevant medical insurance claim records. This happens when the similarity threshold is set too low, or an effective metadata filtering mechanism is lacking, resulting in the recall of data that is semantically similar but irrelevant to the query intent.
How to Verify Configuration
- For typical clinical trial pre-screening questions, such as "Does a patient meet the medical insurance reimbursement criteria for a certain disease?", verify that the recall results include all relevant diagnosis codes, drug codes, and expense details.
- Simulate high-concurrency query scenarios and monitor system response times. Ensure that the average response time remains below the set performance indicator when
qpsreaches the expected peak. - Randomly select
100queries and manually assess the accuracy and completeness of the recall results. Adjust theSimilarity threshold(Similarity Threshold) based on the evaluation.
The values provided are common starting points and should be measured against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.