Data Characteristics
Medical insurance access and pharmacovigilance data primarily originates from documents published by national and local medical insurance bureaus. These include notices for medical insurance catalog adjustments, drug payment standards, negotiation results, clinical medication guidelines, and pharmacovigilance regulations. Document updates are irregular, occurring several times a year or frequently during specific policy changes. Documents typically follow government official document formats. They contain numerous clauses, detailed rules, drug lists (often in tables), approval process descriptions, and clinical efficacy and adverse reaction data. Field names are highly standardized, such as "Generic Drug Name," "Medical Insurance Payment Scope," "Indications," "Adverse Event," and "Reporting Entity." These often include clear units and standard terminology, such as dosage units (mg, g), frequency (times/day), and time periods (years, months).
Constraints on Document Parsing and Chunking
The characteristics of medical insurance access and pharmacovigilance documents impose specific requirements on document parsing and chunking. Irregular updates demand efficient incremental update and version management capabilities to ensure knowledge base timeliness. The coexistence of official document structures and tabular data means a single text chunking strategy is insufficient. Different parsing methods are necessary for structured content (like tables) and unstructured content (like clause descriptions). Standardized fields and units in drug lists require effective identification and preservation of this critical information during chunking. This prevents context loss due to over-chunking. For example, core information like drug names, indications, and adverse reaction types should remain within the same chunk if possible. Additionally, if images (e.g., drug leaflet screenshots, flowcharts) contain critical information, OCR technology is needed to recognize and include their content in the chunking scope.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk Length | 500–800 characters | Medical insurance clauses and adverse reaction descriptions often contain multiple qualifying conditions. This length helps maintain the complete semantic integrity of a single clause. |
Overlap Length | 50–100 characters | Ensures contextual continuity, especially at clause transitions, preventing critical information from being cut off. |
Parsing Strategy | Prioritize by title, combined with fixed-length chunking | Official medical insurance documents often have clear section titles. Chunking by title effectively maintains structural integrity. Fixed length serves as a supplement for continuous text without titles. |
Table Parsing Mode | Structured extraction and conversion to Markdown table | Tabular data like drug lists should retain their structured properties for easier retrieval and display. |
Image OCR | Enabled, prioritize extraction of key labels and text | Some flowcharts or leaflet screenshots may contain critical information regarding medical insurance payment scope or adverse reactions. |
PARSE_FILE_TIMEOUT_SECONDS | 300–600 seconds | Parsing large medical insurance catalog documents can be time-consuming. Increasing the timeout prevents parsing interruptions. |
Common Pitfalls
- System error "Out of VRAM" when uploading large PDF documents: This occurs because the PDF parsing service (e.g., PDF-marker) has insufficient GPU resources or the document is too large, exceeding VRAM limits during processing.
- Some tables in knowledge base documents are not correctly chunked, leading to content disorganization: This often happens when the parser fails to correctly identify complex table structures, or high text density within tables causes the chunking algorithm to fail.
- Images appear abnormal or are missing after importing Word documents: This usually occurs because image paths are not correctly processed into accessible public domain URLs during Word to Markdown conversion, preventing images from loading.
Verification Steps
- Randomly select multiple medical insurance access and pharmacovigilance documents from different sources and formats. Upload them to the knowledge base. Check the chunk preview results, focusing on the completeness of key clauses, drug tables, and adverse reaction descriptions.
- For documents containing complex tables, verify that table content is correctly parsed and converted into readable structured text. Check the correspondence between fields and values.
- Use the knowledge base search function to retrieve specific drug names, indications, or adverse events from documents. Evaluate whether recall results include complete relevant contextual information. Adjust
Recall CountandSimilarity Thresholdbased on actual retrieval performance. - Check the knowledge base version management feature. Confirm that newly added or updated documents correctly generate new versions and that knowledge points in older versions remain unaffected.
Note: The values provided are common starting points. Measure them against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.