Data Characteristics
CAR-T cell therapy clinical trial data originates from national clinical trial registries (e.g., ClinicalTrials.gov, Chinese Clinical Trial Registry) and internal databases of research institutions. This data updates frequently, especially during key milestones like trial recruitment and results publication. Data documents typically include trial protocols, investigator brochures, informed consent forms, and CRF forms. Core fields cover subject inclusion/exclusion criteria, CAR-T product type, target, dose, treatment cycle, adverse events (AEs), serious adverse events (SAEs), and efficacy indicators (e.g., complete response rate CR, partial response rate PR). Field units vary. For instance, dose is measured in cells/kg or cells, treatment cycle in days or weeks, and adverse events often use CTCAE grading.
Constraints from Data Characteristics on Tool Calling and Plugins
The multi-source nature and high update frequency of CAR-T cell therapy clinical trial data require tool calling to have efficient data synchronization and real-time processing capabilities. This ensures pre-screening results are based on the latest information. The complex structure of data documents, especially inclusion/exclusion criteria often described in natural language, demands advanced Natural Language Processing (NLP) capabilities for accurate parsing and conversion into structured query conditions. Diverse field units, such as dose unit conversions, necessitate that tools perform necessary unit standardization or conversion when calling external services. Furthermore, the presence of sensitive information like adverse events imposes higher requirements for data anonymization and privacy protection, impacting data transmission and storage configurations.
Configuration Settings
| Configuration Item | Suggested Value | Rationale |
|---|---|---|
maxContext | 20000 tokens | Accommodates the context requirements of long texts like clinical trial protocols, reducing truncation. |
UPLOAD_FILE_MAX_SIZE | 200 MB | Covers the upload needs for large PDF or compressed clinical trial documents. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Handles parsing complex PDF documents, preventing failures due to parsing timeouts. |
Recall Count | Top 20 | Increases the coverage of relevant trial protocols recalled from the knowledge base, ensuring comprehensive screening. |
Similarity Threshold | 0.75 | Balances recall accuracy and breadth, ensuring screening criteria highly match trial protocols. |
Rerank Return Count | Top 5 | Selects the most relevant trial protocols for further analysis, focusing on key information. |
Common Pitfalls
- Encountering an
HTTP 413 Payload Too Largeerror when calling external APIs typically indicates the request body sent is too large. This exceeds server or gateway limits and requires checking data chunking or compression configurations. - The model fails to correctly identify or apply complex inclusion/exclusion logic during tool calling, leading to inaccurate screening results. This might be due to unclear descriptions of the logic in the prompt or a lack of examples.
- Data synchronization plugins time out, failing to retrieve the latest clinical trial updates. This is usually caused by slow external data source responses or network latency, requiring adjustment of
PARSE_FILE_TIMEOUT_SECONDSor adding a retry mechanism.
How to Verify Configuration
- Simulate multiple clinical trial documents with varying lengths and complexities to verify file upload and parsing functionality. Check parsing logs for errors.
- Select several cases with known inclusion/exclusion criteria. Run the pre-screening process and compare the model's recommended results against expectations. Check the accuracy of field matching.
- Monitor tool call logs. Observe the frequency of external API calls, response times, and return status codes to ensure no abnormal errors or timeouts occur.
Note: The values provided are common starting points. Measure them against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.