Data Characteristics
Medical insurance access quality documents primarily include clinical trial reports, pharmaceutical research data, non-clinical research data, manufacturing process validation files, drug instructions, drug registration certificates, medical insurance payment standard documents, and local medical insurance catalog adjustment notices. These documents often exist in formats such as PDF, Word, and Excel, containing both structured and unstructured content.
Regarding update frequency, medical insurance payment standards and catalog adjustment notices are periodic, typically updated annually or semi-annually. Drug registration certificates and instructions are relatively stable throughout a drug's lifecycle, updating only for significant changes. Documents contain extensive specialized terminology, drug names, disease codes (e.g., ICD-10), payment scopes, reimbursement ratios, and dosage units (mg/kg, U/ml), requiring high precision.
Constraints on Knowledge Base Retrieval and Recall
The periodic updates and high precision requirements of medical insurance access documents necessitate efficient document version management in the knowledge base to ensure retrieval timeliness. Specialized terminology, drug names, and disease codes in documents require domain-specific optimization for word segmentation and entity recognition to prevent semantic drift or loss of critical information.
The presence of multiple document formats challenges the knowledge base's parsing capabilities, requiring accurate content extraction across formats. Numerical information like medical insurance payment standards and reimbursement ratios demands precise matching and contextual association in retrieval results; keyword matching alone is insufficient. This means chunking strategies and recall algorithm selection must prioritize information completeness and logical coherence of context to support accurate decision-making.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Chunk Length) | 800–1200 characters | Medical insurance document paragraphs often contain complete concepts. This length preserves contextual semantics effectively, preventing fragmentation from overly fine-grained chunking. |
Chunk overlap (Chunk Overlap) | 100–200 characters | Ensures critical information at paragraph boundaries is not lost during chunking, improving retrieval recall. |
Recall count (Recall Count) | Top 5 | Medical insurance decisions typically rely on a few highly relevant core documents. Recalling too many increases the LLM's processing burden. |
Similarity threshold (Similarity Threshold) | Calibrate by measurement | Requires adjustment based on specific business scenarios and corpus characteristics using a test set to ensure high relevance recall. |
Rerank result count (Rerank Count) | 3 | Further refines the most relevant items from the recall results, improving the quality and efficiency of the final answer. |
UPLOAD_FILE_MAX_SIZE | 500 MB | Medical insurance documents may contain numerous charts and scanned images. This ensures large files can be uploaded successfully. |
Common Pitfalls
- Receiving an "Error: database error" when creating a new knowledge base may indicate an incorrect
PG_URLconfiguration for Docker Compose deployment, leading to a failed database connection. - LLM responses containing content outside the knowledge base occur when the
Similarity threshold(Similarity Threshold) is set too low. This recalls many irrelevant or low-quality knowledge chunks, diluting effective information. - Missing or inaccurate key numerical values (e.g., reimbursement ratios, dosages) in retrieval results stem from document parsing failures to correctly identify and extract numerical fields from Excel tables or specific PDF layouts.
Verification Steps
- Upload a batch of typical medical insurance access documents (e.g., the latest medical insurance catalog, a drug instruction manual). Verify all documents are parsed and chunked into the knowledge base successfully, without errors.
- For critical information such as medical insurance payment standards and reimbursement scopes, pose questions and observe the
Similarityscores of the recalled results. Ensure core knowledge chunks havesimilarityvalues above the setSimilarity threshold. - For complex queries in the medical insurance access process, evaluate the LLM's answers generated from the knowledge base. Confirm the accuracy and completeness of cited information sources, key numerical values, and specialized terminology by comparing them with the original documents.
The values given are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.