Data Characteristics
CAR-T cell therapy clinical trial pre-screening data comes from patient medical records, gene sequencing reports, flow cytometry results, imaging reports, and prior treatment history. This data exists as unstructured text, semi-structured tables, and structured numerical values. Patient medical records contain free text such as doctor diagnoses, treatment plans, and follow-up notes, with frequent updates (weekly or even daily). Gene sequencing reports are typically PDF or TXT files, containing extensive gene locus information and variant descriptions. Flow cytometry results are often FCS files or reports, providing quantitative cell population data. Imaging reports primarily consist of DICOM images and their interpretation texts. Fields and units are specialized, including tumor burden (e.g., SUVmax values), lymphocyte counts (e.g., CD3+ cell percentage), and gene mutation types (e.g., TP53 mutation). Units are diverse and highly specialized; for example, cell counts use cells/μL, and gene variant descriptions often follow HGVS nomenclature.
Workflow Orchestration Constraints
The diversity of CAR-T cell therapy data imposes specific workflow orchestration requirements. Unstructured patient medical records require robust text extraction and entity recognition capabilities to accurately capture key medical concepts like disease diagnoses, medication dosages, and adverse events from free text. The complex structure and specialized terminology in gene sequencing reports necessitate highly optimized semantic matching for knowledge base construction and retrieval, ensuring effective use of gene locus information in pre-screening rules. High data update frequency means workflow design must incorporate incremental updates and real-time processing mechanisms to avoid reprocessing historical data and ensure pre-screening results are based on the latest information. The heterogeneity of different data sources, such as mixed structured numerical values and unstructured text, requires workflows to flexibly integrate various data processing modules, including numerical comparison, text similarity calculation, and rule engine evaluation. The specialized nature of fields and units constrains the accuracy of conditional judgment nodes in the workflow, requiring precise definition of numerical ranges and text patterns.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
maxContext | 4000 characters | Ensures AI Q&A nodes can process longer patient history summaries and critical gene locus information without exceeding model limits. |
Chunk size | 800 characters | Suitable for segmenting gene sequencing reports and imaging report text, ensuring each knowledge base segment contains sufficient context for effective retrieval. |
Recall count | 5 entries | In knowledge base search nodes, retrieves relevant information from multiple patient documents or gene databases to cover potential pre-screening conditions. |
Similarity threshold | 0.75 | Increases similarity threshold for medical terms and gene variant descriptions to ensure precision of retrieved information and reduce false positives. |
Timeout | 600 seconds | Provides ample time for text parsing, entity extraction, and knowledge base retrieval when processing complex gene sequencing data or multiple medical records. |
Model Name | gpt-4o-mini or claude-3-haiku-20240307 | Balances accuracy in processing complex medical text with inference speed, especially when multi-turn conversations are needed to clarify patient status. |
Common Mistakes
- AI Q&A nodes receive incomplete text, such as only a partial medical history. This occurs when the upstream text concatenation node outputs text exceeding the
maxContextlimit configured for the AI Q&A node. - Knowledge base search nodes in the workflow fail to retrieve expected results, such as a specific gene mutation. This can happen if the knowledge base ID is not correctly passed as a parameter to the node, leading to an incorrect search scope.
- When calling external models for streaming output, the output content is interrupted or has an abnormal format. This indicates that the
res.sendStreammethod in thecodeexecution module was not called correctly or data stream handling was improper, resulting in an incomplete response.
Verification Steps
- Simulate multiple sets of patient data with different pre-screening conditions. Observe if the workflow's final pre-screening results match expectations, especially for borderline cases.
- Inspect the
payloaddata of the output at each critical workflow node (e.g., text extraction, knowledge base search, conditional judgment). Confirm that the data format, field values, and completeness meet expectations. - Use FastGPT's debugging feature to step through the workflow. Check the input and output of each node, particularly focusing on whether the
aiconversation node accurately references medical facts from the knowledge base in response to specific queries. - For complex documents like gene sequencing reports, verify that the knowledge base search node retrieves precise and highly relevant knowledge segments when given specific gene loci or variant descriptions. Adjust
Similarity thresholdto optimize this.
Note: The values provided are common starting points. Measure against your own samples for optimal performance.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.