Data Characteristics
CAR-T cell therapy product data originates from clinical trial reports, drug labels, regulatory approval documents, academic papers, and internal pharmaceutical company R&D documents. This data updates infrequently, typically with clinical trial progress, new product batches, or regulatory policy changes. Document structures are primarily unstructured text, containing extensive medical terminology, experimental data, charts, and references. Core fields include target information, genetic engineering strategies, clinical indications, dosage, administration protocols, side effects, efficacy evaluation metrics (e.g., complete response rate CR, objective response rate ORR), and manufacturing processes. Units commonly used are cells/kg or cells/m² for dosage, days, weeks, or months for time, and CTCAE standards for side effect grading.
Constraints on HTTP Interface and External Systems
The low update frequency of CAR-T cell therapy product data means HTTP interfaces do not require real-time synchronization. A periodic synchronization strategy is suitable, such as full or incremental updates every 24 hours or weekly. Unstructured text, complex medical terminology, and charts require the HTTP interface to handle large file uploads and support document preprocessing services to convert formats like PDF and DOCX into indexable text. The specificity of fields and the use of standards like CTCAE necessitate custom mapping and validation when external systems parse JSON or XML responses to ensure data accuracy. Additionally, potentially large data volumes require robust API concurrency handling and appropriate timeout settings to prevent connection interruptions due from large single requests.
Configuration Settings
| Configuration Item | Suggested Value | Rationale |
|---|---|---|
UPLOAD_FILE_MAX_SIZE | 500 MB | Clinical trial reports and drug labels often contain charts, leading to large file sizes. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Parsing and text extraction from large PDF or DOCX documents can take significant time. |
maxContext | 2000 characters | Ensures capture of complete medical concepts and context, preventing semantic fragmentation. |
Chunk size (Segment Length) | 800–1200 characters | Balances paragraph integrity with model processing efficiency, adapting to medical text paragraph lengths. |
Similarity threshold (Similarity Threshold) | 0.75 | Ensures highly relevant results for professional CAR-T cell therapy consultations. |
HTTP_REQUEST_TIMEOUT | 180 seconds | Provides sufficient time for external systems to process complex queries or large data responses. |
Common Pitfalls
APIinterface returns504 Gateway Timeouterror: This typically occurs when the external system takes too long to process a request, and FastGPT'sHTTP_REQUEST_TIMEOUTparameter is set too short, failing to wait for the external system's response.- Missing critical medical terms or dosage units in query results: This happens when the
JSONstructure returned by the external system is not fully mapped or parsed, and specific fields (e.g.,CTCAEgrades,cells/kg) are not correctly extracted. - FastGPT cannot retrieve the workflow's opening statement: This is due to FastGPT's interface design, where the workflow's opening statement is not directly exposed via an independent
APIinterface. It requires retrieval through specificAPIendpoints or configuration parameters.
Verification Steps
- Upload a CAR-T cell therapy
PDFdocument larger than100 MB. Verify successful parsing and knowledge base chunk generation. - Use the FastGPT interface to query an external system for clinical indications of a specific CAR-T product. Verify that key efficacy metrics like
CRandORRare complete in the returned results. - Simulate high-concurrency requests (e.g.,
10concurrent requests). Observe if FastGPT'sAPIcalls to the external system remain stable, without timeouts or connection errors. - After external system data updates, verify that the FastGPT knowledge base retrieves the latest data within the configured synchronization cycle (e.g.,
24 hours).
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.