Data Characteristics
Data for rational drug use regulations primarily originates from policies and regulations published by the National Health Commission and the National Medical Products Administration, clinical treatment guidelines, drug inserts, and internal Standard Operating Procedures (SOPs) developed by medical institutions. These documents typically exist in PDF, Word, or structured text formats. Data updates are frequent; drug inserts and clinical guidelines, in particular, may be revised quarterly or annually. Document structures are complex, containing extensive professional terminology, dosage instructions, contraindications, drug interactions, and often include nested tables and diagrams. Core fields include drug name, generic name, indications, usage and dosage, adverse reactions, guidance for special populations, drug interactions, storage conditions, and relevant regulatory clause numbers. Units involved include milligrams (mg), milliliters (ml), times/day, and treatment duration (days).
Constraints Imposed by These Characteristics on "HTTP API and External Systems"
The authoritative nature and high update frequency of rational drug use data sources require HTTP APIs to have efficient file upload and parsing capabilities, along with support for version control. Complex document structures and specialized terminology mean traditional keyword matching often misses critical information, necessitating vector retrieval combined with semantic understanding. The presence of nested tables and diagrams challenges the structural extraction capabilities of document parsers, potentially leading to the loss or misplacement of key information. The need for standardized fields and units demands strict preprocessing and validation before data ingestion to ensure accuracy in subsequent Q&A. Furthermore, internal SOPs from different medical institutions may vary in format, requiring the API to be flexible enough to accommodate diverse data inputs.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
chunkSize | 500-800 characters | Balances semantic completeness with retrieval efficiency, avoiding excessively large or small information blocks. |
overlapSize | 100 characters | Ensures contextual continuity and reduces semantic breaks caused by chunking. |
embeddingModel | text-embedding-ada-002 or equivalent model | Addresses the need to understand specialized terminology and complex semantics. |
maxRequestTimeout | 600 seconds | Accommodates the time required for parsing and uploading large PDF or Word files. |
maxRetryAttempts | 3 times | Handles intermittent network fluctuations or temporary service unavailability of external systems. |
dataValidationSchema | Defined by actual document fields | Ensures the structured nature and accuracy of imported data, e.g., drug names, dosage units. |
Common Pitfalls
- After batch importing data, unique identifiers for each data entry are unavailable, making local mapping difficult. This occurs because the API design does not return a detailed list of import results.
- A large amount of data in the knowledge base has been indexed, but search test results show inaccurate recall. This happens because the chunking strategy or embedding model fails to effectively capture the specialized semantics within rational drug use documents.
- Frequent connection timeouts or service unavailability occur when external systems connect to FastGPT. This is due to
maxRequestTimeoutbeing set too low, failing to adequately account for the time required for large data transfers or complex processing.
Verification Steps
- Upload a batch of rational drug use regulation documents in various formats (PDF, Word, TXT). Check if document parsing is successful, if content is fully imported, and if an ID can be obtained for each imported data entry.
- Perform search tests for core rational drug use questions (e.g., "contraindications of drug XX," "interaction between drug XX and drug YY"). Observe the accuracy and relevance of recall results and evaluate their match with expected answers.
- Simulate data addition and update operations from an external system via the HTTP API. Check if the API returns a success status code (e.g.,
200) and verify the actual changes in the knowledge base. - Review FastGPT logs for any error messages during data import and querying, especially those related to data parsing, vectorization, or external system communication.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.