Data Characteristics for This Category
Respiratory system quality documents typically include clinical guidelines, diagnostic standards, treatment plans, drug inserts, and case reports. These documents originate from various sources, such as authoritative medical organizations, pharmaceutical companies, internal hospital regulations, and clinical trial data. Update frequency is relatively high, especially with new drug releases, treatment breakthroughs, or guideline revisions. Some core documents may update every six months to a year, while clinical trial data might release quarterly. Document structures often use PDF, DOCX, or RTF formats, frequently containing hierarchical headings, charts, flowcharts, and specialized terminology. Fields involve drug dosages (e.g., mg/kg), treatment durations (e.g., days, weeks), disease staging (e.g., GOLD staging), and laboratory indicators (e.g., FEV1, SpO2), with precise and standardized units.
Constraints Imposed by These Characteristics on Tool Calling and Plugins
The update frequency and diverse data sources of respiratory system quality documents require tool calling and plugins to have flexible data synchronization and integration capabilities. The large volume of specialized terminology and measurement units in documents demands high accuracy in text parsing. Optimization for medical vocabulary is necessary to prevent information loss or misinterpretation due to improper lexical analysis. Flowcharts and tabular data are often overlooked in traditional text parsing, limiting the effectiveness of text-only retrieval. Plugins must identify and extract structured information. For example, drug dosage calculations or treatment plan comparisons often require precise extraction of numerical values and units from tables. Furthermore, if classification information like disease staging is not effectively extracted and linked to a knowledge graph, it impacts subsequent rule-based decision support. This poses a challenge to the semantic understanding and entity recognition capabilities of plugins.
Configuration Recommendations
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
maxContext | 3000 | Respiratory system documents are often lengthy, requiring a larger context window to capture complete information. |
Chunk size (Segment Length) | 800–1200 characters (characters) | Balances semantic completeness and retrieval efficiency, avoiding over-segmentation that breaks context. |
Similarity threshold (Similarity Threshold) | 0.75–0.85 | Medical texts are highly specialized; a high threshold helps exclude low-relevance results and improves accuracy. |
Recall count (Retrieval Count) | Top 5 entries (top 5) | Ensures concise retrieval results, reduces interference from irrelevant information, and focuses on core content. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds (seconds) | Large PDF or DOCX documents take longer to parse; this reserves sufficient parsing time. |
ENABLE_STRUCTURED_PARSING | true | Respiratory system documents contain many tables and flowcharts; enabling structured parsing extracts key data. |
Three Common Mistakes
- Receiving an
HTTP 400 Bad Requesterror when calling external APIs. A common reason is that JSON field names or data types in the request body do not match the API documentation, for instance, writingpatientIdinstead ofpatient_id. - Key fields (e.g.,
drug_dosageortreatment_duration) are empty in the results returned after plugin execution. This may occur if the document parsing did not correctly identify or extract numerical values and units from tables. - The system encounters an
OutOfMemoryErroror parsing timeout when processing PDF documents with many flowcharts. This typically happens ifPARSE_FILE_TIMEOUT_SECONDSis set too low, or the parser cannot effectively handle complex graphical elements.
How to Verify Configuration
- Upload a respiratory system clinical guideline PDF containing complex tables and flowcharts. Check if the parsed text and structured data are complete and accurate, especially for critical information like drug dosages and treatment plans.
- Use FastGPT's knowledge base log feature to review parsing logs for specific documents. Confirm if any parsing failures or warnings exist, and verify if
PARSE_FILE_TIMEOUT_SECONDSis sufficient. - Test the knowledge base's retrieval results with queries containing specific medical terms and measurement units. Evaluate the effect of
Similarity threshold(Similarity Threshold) andRecall count(Retrieval Count) to ensure high relevance and an appropriate number of results. - Build a simple workflow that calls an external drug database API. Input a respiratory system disease drug name and verify if the plugin correctly passes parameters and retrieves drug inserts or interaction information.
Note: The values provided are common starting points. It is recommended to measure them against your own samples for optimal performance.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.