Data Characteristics
Medical insurance settlement drug vigilance data originates from settlement manifests of medical insurance agencies, drug procurement records, and patient medication details uploaded by some medical institutions. Data updates frequently, typically in monthly or quarterly batches, involving historical data backfilling and incremental synchronization. The data structure is primarily structured tables, commonly in CSV, XML, or JSON formats. Key fields include patient unique identifier, generic drug name, dosage form, specification, dose, administration route, medical insurance payment category, settlement amount, settlement date, medical institution code, and diagnosis information (ICD-10 code). Units for dose are often milligrams (mg), grams (g), or milliliters (ml); amounts are in CNY; timestamps are precise to the day.
Constraints on Deployment and Upgrade
The batch update nature of medical insurance settlement data requires deployment solutions to support efficient batch data import and incremental update mechanisms, avoiding frequent small file operations. The predominance of structured data means the knowledge base vectorization process must focus on accurate field mapping and structured information extraction. The specialized and standardized nature of fields like medical insurance payment categories and diagnosis codes demands higher accuracy in text segmentation and recall strategies, ensuring correct identification of relevant terminology. Data volume is typically large, especially in national or provincial deployments, posing challenges to server storage capacity, memory, and I/O performance, potentially leading to read/write bottlenecks during knowledge base construction or querying. The data update frequency dictates the scheduling cycle and resource reservation for timed tasks.
Configuration Recommendations
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
UPLOAD_FILE_MAX_SIZE | 2048 MB | Medical insurance settlement data files can be large, especially during monthly or quarterly batch updates, requiring support for uploading large files. |
Chunk size (Segment Length) | 800–1200 characters (characters) | Medical insurance settlement records typically contain multiple structured fields. Longer segment lengths help preserve the complete context of a single record, preventing truncation of key information. |
Recall count (Recall Count) | Top 10 entries (top 10 items) | For queries on specific drugs or diagnoses in medical insurance settlements, increasing the recall count can improve coverage of relevant records, especially when similar drugs or diagnosis codes exist. |
Similarity threshold (Similarity Threshold) | 0.75 | Drug names and diagnosis codes in medical insurance settlement data are highly standardized. Setting a relatively high similarity threshold helps achieve precise matching and reduces false recalls. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds (seconds) | Parsing large medical insurance settlement data files during batch import can take a long time. Extending the timeout prevents parsing interruptions. |
maxContext | 3072 | Medical insurance settlement data involves a large amount of field information. Increasing the context length helps the language model understand complex relationships, such as those between drugs, diagnoses, and payment categories. |
Common Pitfalls
- During knowledge base construction, encountering a
MongoServerError: The dollar ($) perror usually indicates an incompatibility between the MongoDB version and certain FastGPT plugins or features. Check if themongoversion meets the requirements of the current FastGPT version. - After local deployment, when creating knowledge base vectors or retrieving knowledge via the API, frequent server read/write performance bottlenecks manifest as slow system response or freezes. This often results from
UPLOAD_FILE_MAX_SIZEand other parameters being set too low, or insufficient server disk I/O and memory configuration to handle the high-concurrency read/write demands of medical insurance settlement data import and vectorization. - After exporting a knowledge base backup and importing it into another FastGPT instance of the same version, if the number of imported knowledge base files does not match the exported count, this may relate to differences in the backup export mechanism when handling multi-source or multi-type knowledge bases. Verify the
FastGPTV4.13.0version's backup logic to ensure all files are correctly packaged.
Verification Steps
- Perform a batch import test with medical insurance settlement data. Observe system logs to confirm file parsing and vectorization processes complete without errors. Check if
PARSE_FILE_TIMEOUT_SECONDSsettings cover the import duration. - Conduct test queries using key fields such as medical insurance payment categories, generic drug names, and diagnosis codes. Verify if the
Recall count(Recall Count) andSimilarity threshold(Similarity Threshold) of the returned results meet expectations, ensuring relevant information is accurately recalled. - Monitor server CPU, memory, and disk I/O utilization. During knowledge base construction or high-concurrency retrieval tasks, ensure resource utilization remains within acceptable limits, without sustained high load or resource exhaustion. Specifically, verify that large file uploads corresponding to
UPLOAD_FILE_MAX_SIZEare stable.
Note: The values provided are common starting points. Measure against specific samples to determine optimal settings.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.