Data Characteristics
Medical insurance access drug safety data primarily originates from national and local medical insurance catalogs, drug specifications, clinical study reports, real-world evidence (RWE), and adverse drug reaction (ADR) monitoring databases. This data updates frequently. Medical insurance catalogs typically adjust annually, with some provinces and cities making quarterly or semi-annual dynamic adjustments. Drug specifications and clinical trial data continuously update with approval progress and post-market studies. ADR data is collected continuously.
Regarding document structure, medical insurance catalog data is often published in structured table formats, including fields like drug codes, generic names, dosage forms, specifications, and payment scope. Drug specifications and clinical reports are primarily unstructured text, covering dosage, usage, contraindications, and adverse reaction event descriptions. Specific fields include "medical insurance payment scope" and "restricted payment conditions." Units often involve "yuan/unit," "times/day," and "mg/kg."
Constraints on Deployment and Upgrade
The update frequency and heterogeneity of medical insurance access data impose specific requirements on FastGPT deployment. Regular updates to medical insurance catalogs mean the knowledge base must support efficient bulk import and incremental synchronization, and handle differences between various catalog versions. The large volume of unstructured text (e.g., drug specifications) requires more refined text segmentation and entity extraction during data preprocessing to ensure accurate knowledge retrieval. The presence of specific fields, such as payment conditions, necessitates customized parsing logic to prevent loss or misinterpretation of critical information.
Furthermore, continuous input of ADR data requires a deployment solution with stable data ingestion channels and real-time knowledge update capabilities. This ensures the AI Agent can make decisions based on the latest information. Issues with variables not returning during API calls may relate to the model's context window not correctly loading the latest variable definitions after data source updates.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
UPLOAD_FILE_MAX_SIZE | 500 MB | Medical insurance catalogs or clinical study reports often contain large amounts of text and tables, leading to larger file sizes. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Processing large PDF or Excel format medical insurance catalog files can take a long time. |
Chunk size (Segment Length) | 800–1200 characters | Drug specification and clinical report paragraphs are long; overly short segments risk losing context, while overly long ones increase irrelevant information. |
Recall count (Recall Count) | Top 10 entries | Medical insurance payment conditions and adverse reaction event descriptions may be scattered across multiple related paragraphs; increasing recall count improves coverage. |
Similarity threshold (Similarity Threshold) | 0.78–0.82 | Precisely match medical insurance policies and drug characteristics to avoid misjudgments due to semantically similar but critically different information. |
Rerank result count (Reranked Return Count) | Top 5 entries | After initial recall, reranking further filters for the most relevant payment conditions or adverse reaction cases related to the query. |
Common Mistakes
- Symptom: Some variable data does not return during API calls, or the returned results differ from debugging. Reason: The knowledge base data source updated after deployment, but the variables referenced in the workflow or their corresponding knowledge snippets were not synchronized or re-indexed in time.
- Symptom: Uploading large medical insurance catalog files leads to a long unresponsive interface or a
504 Gateway Timeouterror. Reason: ThePARSE_FILE_TIMEOUT_SECONDSparameter is set too low, insufficient to process documents containing complex table structures or large amounts of text. - Symptom: Medical insurance policy query results are inaccurate, often providing outdated or inapplicable payment scopes. Reason: The knowledge base does not integrate with regular update mechanisms for medical insurance catalogs, or incremental updates fail to correctly identify and replace old version data.
Verification Steps
- Select a recently updated medical insurance catalog file, upload it, and observe its parsing status. Confirm that all table data and text content are correctly extracted.
- For a specific drug, pose a query including "medical insurance payment scope" or "restricted payment conditions." Compare the returned results with the original medical insurance catalog document to ensure critical field information is accurate.
- Simulate an adverse drug reaction event report. Test whether the Agent can recall relevant adverse drug reaction cases from the knowledge base based on symptom descriptions. Check if the recall count and similarity meet expectations.
The values given are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.