Data Characteristics
Medical insurance access pharmacovigilance data originates from various review materials submitted by Marketing Authorization Holders (MAH), policy documents from the National Healthcare Security Administration, local medical insurance catalog adjustment announcements, and pharmacoeconomic evaluation reports. Data updates are relatively fixed, typically aligning with national or local medical insurance catalog adjustment cycles, which can be semi-annual or annual. Document structures are primarily structured and semi-structured, including drug instructions, clinical trial reports, pharmacoeconomic model data, and medical insurance payment standard documents. Fields include generic drug name, indications, dosage and administration, medical insurance coverage, reimbursement ratio, adverse event codes (e.g., MedDRA codes), incidence, severity, causality assessment, and cost-effectiveness ratios (ICER values) in economic evaluations. Units commonly used are milligrams (mg), grams (g) for dosage; times/day, times/week for frequency; RMB yuan (RMB) for cost; and days, months, years for time.
Constraints on Vector Models and Indexing
The periodic update cycle of medical insurance access data dictates the vector index rebuilding or incremental update strategy. This strategy must align with medical insurance policy release schedules to prevent data lag from affecting decisions. The mix of structured and semi-structured data requires vector models to effectively process tables, text descriptions, and numerical information. For example, ICER values in pharmacoeconomic reports need embedding alongside text descriptions to maintain semantic integrity. MedDRA codes for adverse events are standard terminology; their hierarchical relationships and conceptual associations must be preserved during vectorization to avoid semantic loss from simple string matching. The precision requirements for key fields like medical insurance coverage and reimbursement ratios necessitate high-precision similarity matching during retrieval. This may also require rule-based post-processing mechanisms to ensure critical information is not over-generalized. Integrating and deduplicating data from different sources adds complexity to index construction, requiring consideration of document ID uniqueness and version management.
Configuration Settings
| Configuration Item | Suggested Value | Rationale |
|---|---|---|
Chunk size | 500-800 characters | Balances long text semantic integrity with short text recall efficiency |
Chunk overlap | 50-100 characters | Ensures context continuity, reduces information loss |
Recall count | 8-12 entries | Covers potentially relevant information, balances computational resources and recall quality |
Similarity threshold | Calibrate by measurement, e.g., 0.75-0.85 | Ensures highly relevant results, reduces noise interference |
Rerank result count | 3-5 entries | Refines final results, improves user experience |
maxContext | 4096-8192 token | Accommodates complex queries and long text contexts, prevents truncation |
Common Pitfalls
- Irrelevant drug or payment scope information appears in query results. This occurs when the vector model insufficiently distinguishes between generic and brand names, or inaccurately interprets medical insurance payment scope conditions.
- Query results remain on old version information after data updates. This occurs when the index is not rebuilt promptly or incremental updates fail, causing the vector database to be out of sync with the latest medical insurance policies.
- Deviations in medical insurance reimbursement ratios or restriction conditions occur. This happens when numerical fields are not effectively integrated with corresponding text descriptions during vectorization, leading to semantic information loss.
Validation Steps
- Select recently released medical insurance policy documents. Query their core content and verify that the returned results include key drug names, payment scopes, and reimbursement conditions.
- Randomly select a batch of drug instructions containing adverse event information. Query the incidence and severity of specific adverse reactions and check the accuracy of the returned results.
- Compare the number of retrieved items for medical insurance access-related queries against the expected range. Observe changes in precision and recall after adjusting the
Similarity thresholdto determine an appropriate threshold.
The values provided are common starting points. Measure them against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.