Data Characteristics for This Category
CAR-T cell therapy product data originates from diverse sources, including clinical trial reports, drug labels, research papers, regulatory databases (e.g., FDA, EMA), and patent literature. Data update frequencies vary; clinical trial data typically updates with research progress or results, drug labels change with version iterations, and research papers are continuously published. Document structures are highly specialized, containing detailed pharmacological, pharmacokinetic, clinical efficacy, and safety data. Specific fields include target antigens, transduction vector types, cell manufacturing processes, dosages (e.g., cells/kg), administration regimens, and adverse event rates (e.g., CRS grade, ICANS incidence). Units strictly adhere to biomedical standards, such as 10^6 cells/kg, mg/kg, %, and days.
Constraints on "Forms and Interaction" Imposed by These Characteristics
The specialized and complex nature of CAR-T cell therapy product data imposes specific requirements on form and interaction design. Data dispersion means the system needs to support importing and integrating multi-source heterogeneous data and handle various document formats (e.g., PDF, HTML, XML). Uncertain update frequencies require the RAG knowledge base to have flexible data synchronization mechanisms, capable of periodically checking and updating relevant data. The complexity of document structures necessitates precise parsing of nested structures, tabular data, and unstructured text during information extraction, such as accurately identifying primary endpoints, secondary endpoints, and adverse event data from clinical trial reports. The strictness of fields and units requires form input to provide precise unit prompts and validation rules to prevent data confusion due to incorrect units. Furthermore, the presence of numerous specialized terms means form interaction design needs to support intelligent term suggestions and explanations to reduce user input barriers.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
UPLOAD_FILE_MAX_SIZE | 500 MB | Clinical trial reports and drug labels often contain numerous charts and detailed text, resulting in large file sizes, requiring support for large file uploads. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Processing complex PDF documents, especially those containing many tables and scanned images, takes a long time. |
Chunk size (Segment Length) | 800 characters | Key information density in CAR-T product descriptions is high; overly short segments may break context, while overly long segments reduce recall accuracy. |
Recall count (Recall Count) | Top 8 | Ensures coverage of multi-dimensional relevant information in complex queries, such as efficacy, safety, and indications. |
Similarity threshold (Similarity Threshold) | 0.75 | Ensures recalled segments are highly relevant to the user's query, avoiding the introduction of significant noise that could affect the accuracy of CAR-T product assessments. |
Entity Recognition Model | Proprietary Model | General models lack sufficient recall accuracy for specialized entities like CAR-T targets, cell types, and adverse events, requiring specific training. |
Three Common Pitfalls
- The system fails to recognize specialized terms entered by the user or provides incorrect prompts. This is because the RAG knowledge base lacks a professional vocabulary and entity recognition capabilities specifically for the biomedical field, particularly CAR-T therapy.
- Query results for CAR-T product dosages or adverse event rates show unit confusion or omission. This is due to a lack of standardization of field units from different sources during data ingestion, or the parser failing to correctly extract unit information.
- In multi-turn conversations, the AI cannot consistently ask follow-up questions about specific CAR-T product manufacturing process details based on context. This is because the knowledge base segmentation strategy is too coarse, leading to related information being truncated or dispersed, lacking coherence.
How to Verify Configuration
- Select 5 CAR-T product labels or clinical reports from different sources, upload them to the system, and check if file parsing is complete and if key fields (e.g.,
target,indication,adverse events) are correctly identified and extracted. - Construct 10 complex queries containing specialized terms, such as "What is the
ORRof a specific CAR-T product for a particular lymphoma?", and verify the accuracy and completeness of the AI's answer, as well as the correctness of cited sources. - Simulate a user engaging in multi-turn conversations, asking in-depth questions about a CAR-T product's manufacturing process, dosage adjustments, or safety management, to assess whether the AI can maintain contextual consistency and provide coherent, professional answers.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.