Data Characteristics
CAR-T cell therapy product data originates from clinical trial reports, drug labels, regulatory approval documents, academic papers, and internal pharmaceutical R&D documents. Data updates are infrequent, typically occurring quarterly or annually, aligned with clinical progress or regulatory cycles. Documents are primarily structured PDFs, containing extensive medical terminology, dosage units, clinical indicators, and adverse event descriptions. Fields and units are highly specialized. For example, "dosage" often appears as cells/kg, "efficacy" involves standards like CR (Complete Remission) and PR (Partial Remission), and adverse events focus on CTCAE grades.
Constraints on Model Integration and Configuration
The specialized nature of CAR-T cell therapy data imposes specific requirements on model integration and configuration. Infrequent updates mean initial data ingestion may be time-consuming, but subsequent incremental updates are less demanding. Therefore, focus on the initial setting of the PARSE_FILE_TIMEOUT_SECONDS parameter. Documents contain complex medical terms and specialized abbreviations. The model must handle these effectively during tokenization and entity recognition, requiring appropriate dictionaries or pre-trained models. Table and chart information within structured text can be lost during standard text extraction, necessitating enhanced file parsing capabilities. Standardized fields and units require rigorous unit normalization during data preprocessing to prevent discrepancies in model understanding. For example, 10^6 cells/kg and million cells/kg from different sources need unification.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
UPLOAD_FILE_MAX_SIZE | 500 MB | Clinical reports and drug labels are large; ensure single-upload capability. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Complex PDF parsing is time-consuming; prevent timeout failures. |
Chunk size (Chunk Length) | 800–1200 characters | Balances context completeness and model processing capacity for medical texts. |
Recall count (Recall Count) | Top 8 entries (Top 8) | Ensures coverage of key information from multiple relevant documents, improving answer accuracy. |
Similarity threshold (Similarity Threshold) | Calibrate based on actual measurements | Adjust based on the semantic similarity distribution of the specific dataset to avoid false positives or negatives. |
Rerank result count (Reranked Return Count) | Top 3 entries (Top 3) | Focuses on the three most relevant pieces of information, reducing model processing load and improving timeliness. |
Common Pitfalls
- Model output for dosage or efficacy data shows inconsistent units or numerical errors. This occurs when medical units are not standardized during data preprocessing.
- The model fails to accurately extract adverse event rates from clinical trial reports. Related fields are empty or information is missing. This happens when the file parser does not effectively identify and extract data from complex table structures.
- When queried about specific medical terms, the model provides generic answers or indicates "I cannot provide relevant information." This occurs when the model lacks integration with a specialized medical dictionary or relevant domain knowledge during training.
Verification Steps
- Upload a typical CAR-T clinical trial report containing dosage, efficacy, and adverse events. Check if the model accurately extracts and understands this key information, especially descriptions involving
CTCAEgrades. - Query specific parameters for the same CAR-T product from different sources (e.g., drug labels and academic papers). Verify if the model integrates information and provides consistent answers.
- Input complex questions with specialized medical terminology. Observe the professionalism and accuracy of the model's response to determine its effectiveness in handling domain-specific vocabulary.
Note: The values provided are common starting points. Measure against your own samples for optimal performance.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.