Data Characteristics for This Category
Contraindication and interaction data primarily originates from drug inserts, pharmacopoeias, clinical guidelines, drug databases (e.g., Lexicomp, Micromedex), and professional literature. The update frequency is relatively stable, typically released quarterly or annually, coinciding with drug approval updates, new drug launches, or clinical research advancements. The document structure is mainly structured or semi-structured data, often in JSON, XML format, or database records. Each entry includes fields such as drug name, generic name, active ingredient, interacting drugs, interaction type (e.g., pharmacodynamic, pharmacokinetic), clinical manifestations, severity, and management recommendations. Units for dosage are commonly milligrams (mg), grams (g), or milliliters (ml), and time in hours (h) or days (d). However, interaction descriptions are predominantly textual and qualitative.
Constraints on Deployment and Upgrade from These Characteristics
The structured nature of contraindication and interaction data allows FastGPT to perform document parsing and chunking more efficiently during knowledge base construction. The periodic data updates require reserving automated or semi-automated data synchronization and knowledge base reconstruction mechanisms during deployment. Since interaction descriptions are often long and contain specialized terminology, knowledge base chunking must avoid semantic breaks, ensuring each chunk completely expresses an interaction's logic. Furthermore, interaction severity and management recommendations are critical information; vector retrieval must ensure accurate recall of these core fields. The deployment environment needs to support efficient indexing and retrieval of large amounts of text data, while also requiring atomicity and rollback capabilities for update operations to handle potential errors during data synchronization.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Chunk Length) | 800–1200 characters | Ensures completeness of interaction descriptions, preventing semantic truncation. |
Overlap Length | 100 characters | Guarantees contextual continuity between chunks, improving recall quality. |
Max Chunks | 20 | Limits excessively long single documents, balancing retrieval efficiency and recall precision. |
Recall count (Recall Count) | Top 5–8 | Covers potential related information while avoiding interference from irrelevant content. |
Similarity threshold (Similarity Threshold) | Calibrate based on actual measurements | Requires multiple test calibrations based on specific data and retrieval effectiveness. |
PARSE_FILE_TIMEOUT_SECONDS | 300 seconds | Addresses parsing time when processing large drug inserts or batch data imports. |
Three Common Mistakes
- After importing knowledge base documents, Q&A results lack critical interaction severity or management recommendations: This usually results from excessively fine-grained document chunking, which disperses core information across different chunks, leading to incomplete recall during retrieval.
- After upgrading FastGPT, Docker containers fail to start or some functions are abnormal: Common causes include the
docker-compose.ymlconfiguration not being updated with the new version, or mount volume permission issues preventing containers from writing necessary files. - Contraindication and interaction Q&A results contain a large amount of irrelevant or vague content: This often happens when the
Similarity threshold(Similarity Threshold) is set too low, leading to the recall of knowledge chunks weakly or unrelated to the user's query.
How to Verify Correct Configuration
- Select representative drug interaction queries. Check if the Q&A results accurately mention relevant contraindications, interacting drugs, clinical manifestations, and management recommendations. Compare these with the original data.
- Simulate the data update process. Observe if the knowledge base reconstruction is successful and verify if the new and old data are reflected as expected in the Q&A results.
- In the FastGPT knowledge base management interface, randomly view several chunked contraindication and interaction documents. Verify if the chunk length and content maintain semantic integrity, especially in the interaction description section.
- Use FastGPT's debugging tools to view the recalled chunk content and their
similarityvalues for specific queries. Determine if theRecall count(Recall Count) andSimilarity threshold(Similarity Threshold) settings are appropriate.
The values provided are common starting points. Measure against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.