Data Characteristics
CAR-T cell therapy product data comes from various sources. These include clinical trial databases (e.g., ClinicalTrials.gov, national clinical trial registration platforms), regulatory agency approval documents (e.g., FDA, EMA, NMPA), academic journal literature, and biotechnology company official product manuals. Data update frequencies vary. Clinical trial data changes with recruitment, enrollment, and results publication. Approval documents are relatively stable post-approval but may update with label revisions or indication expansions.
Data document structures are complex. They typically contain large amounts of unstructured text (e.g., clinical study reports, drug mechanism descriptions) and structured data (e.g., indications, dosages, adverse reactions, manufacturers, batch information). Fields cover biomarkers, targets, gene sequences, clinical endpoints, dosing regimens, and pharmacokinetic parameters. Units include concentration (e.g., ng/mL), time (e.g., days), and cell count (e.g., cells/kg). Abbreviations and specialized terminology are common.
Constraints on Tool Calling and Plugins
The complexity of CAR-T cell therapy product data imposes specific constraints on tool calling and plugins. A high proportion of unstructured text requires strong text parsing capabilities. Traditional query plugins based on fixed fields may not effectively extract key information.
Heterogeneous data from multiple sources demands flexible data integration and standardization capabilities from plugins. For example, mapping synonymous terms across different databases. Varying update frequencies, especially for clinical trial progress, require tool calls to trigger periodic or on-demand checks for data source updates to ensure information timeliness.
The use of specialized terminology and abbreviations challenges the semantic understanding and entity recognition capabilities of plugins. This requires enhancement with domain-specific lexicons. Numerical data involving drug dosages and biomarkers requires plugins to handle unit-aware numerical comparisons and calculations, preventing errors from unit mismatches.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
maxContext | 4096 | Adapts to complex document structures, ensuring sufficient context for understanding. |
Chunk size | 500–800 characters | Balances semantic completeness with model processing efficiency, preventing loss of focus from overly long segments. |
Recall count | Top 10–15 entries | Increases the likelihood of recalling relevant information from multiple sources, covering a wider range of potential answers. |
Similarity threshold | 0.75 | Filters out low-relevance results, focusing on CAR-T product-specific queries. |
Rerank result count | 5 | Ranks the most relevant results higher after a high recall volume. |
HTTP_TIMEOUT | 60000 ms | Addresses slow responses from external API calls due to large data volumes or network latency. |
Common Mistakes
- External API calls return
502errors. Common causes include failed authentication for the external service interface or incorrect request parameter format, preventing the gateway from forwarding the request correctly. - Search plugins return web page links that cannot be automatically parsed. This results in only a URL list or summary information. The reason is a lack of further web content crawling and parsing capabilities, preventing the model from directly reading link content.
- Some models in the tool list are not selectable. This may be because the corresponding models are not enabled or correctly configured in the
MODEL_CAPABILITIESsetting, causing them not to appear as available options in the interface.
Verification
- Simulate various CAR-T product consultation scenarios. For example, query specific product targets, indications, and adverse reactions. Check if tool calls accurately return the corresponding information.
- Verify if plugins successfully parse and extract key fields from external web pages regarding CAR-T clinical trial progress or latest approval updates.
- Review system logs to confirm external API calls return a
200status code and that the data structure meets expectations without parsing errors.
The values given are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.