Data Characteristics in this Domain
Respiratory diseases, such as asthma, chronic obstructive pulmonary disease (COPD), and lung cancer, involve clinical trial data from diverse and heterogeneous sources. Data originates from national clinical trial registries (e.g., ClinicalTrials.gov, Chinese Clinical Trial Registry), internal drug development databases, Electronic Health Record (EHR) systems, and various biomarker testing platforms. Data update frequencies vary. Registry information typically updates upon trial initiation, modification, or result publication, while individual patient data may be imported in real-time or periodically. Document structures are complex, including trial protocols, informed consent forms (ICFs), case report forms (CRFs), medical imaging reports (e.g., CT, MRI), genetic sequencing data, and various laboratory test reports. Fields and units are diverse. Examples include lung function indicators (FEV1, FVC, in L or %), blood gas analysis (PaO2, PaCO2, in mmHg), imaging lesion sizes (long diameter, short diameter, in mm), gene mutation sites (e.g., EGFR, ALK), and patient symptom scores (e.g., CAT score).
Constraints Imposed by these Characteristics on Tool Calling and Plugins
The broad range of respiratory clinical trial data sources requires tool calling and plugins to have robust multi-source data integration capabilities. Heterogeneous data formats and varying update frequencies mean that data retrieval must adapt to different API interfaces and consider incremental and full update strategies. Complex document structures and diverse field units demand high standards for information extraction and standardization. Plugins must parse various report formats and perform unit conversions and concept mapping. For example, lung function data may exist as raw values or percentages; plugins need to handle these uniformly. Furthermore, access control and data anonymization for sensitive data (e.g., genetic sequencing) become critical security and compliance constraints for tool calling. The need for specific biomarkers requires plugins to precisely query and extract relevant fields and support complex logical combination queries to meet pre-screening conditions.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
maxContext | 4000 token | Balances information volume from multiple documents with inference efficiency, preventing context overflow. |
PARSE_FILE_TIMEOUT_SECONDS | 300 seconds | Processing large PDF trial protocols or imaging reports can be time-consuming. |
Chunk size | 800 characters | Ensures individual text blocks contain sufficient semantic information, reducing context fragmentation. |
Recall count | Top 10 entries | Covers more potentially relevant documents, improving pre-screening accuracy. |
Similarity threshold | 0.75 | Balances recall and precision, filtering out irrelevant clinical trial information. |
tool_request_timeout | 60 seconds | Accounts for potential network latency or processing time during external API calls. |
Common Mistakes
- When calling an external API, the conversation log displays a title, but the actual data fields are empty. This usually happens because the API's returned data structure does not match the
output_schemadefined in the plugin, leading to parsing failure. - After asking the model several questions, frequent
context window exceedederrors occur. This happens whenmaxContextorChunk sizeare not set appropriately, causing too much input information in a single conversation. - When using FastGPT's API for RAG-based questioning, the expected knowledge base recall results are not obtained. This might be due to
Similarity thresholdbeing set too high, or the knowledge base not being correctly imported with respiratory-related documents.
Verification of Configuration
- Execute a series of pre-screening queries containing respiratory disease characteristics (e.g., asthma, FEV1, EGFR mutation). Check if the results include relevant clinical trial information and verify key field values.
- Simulate calls to external data interfaces. Observe the logs to see if
tool_request_timeoutis triggered and check if the returned JSON data structure strictly matches theoutput_schema. - In the FastGPT interface, review the conversation logs. Confirm that the knowledge base passages cited by the model for pre-screening judgments are accurate and complete, and that no
context window exceededwarnings appear.
The values provided are common starting points and should be measured against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.