Data Characteristics
CAR-T cell therapy pharmacovigilance data originates from clinical trial reports, real-world evidence (RWE) studies, individual case safety reports (ICSRs), and regulatory databases. This data updates frequently, sometimes weekly or even daily, especially during new product launches. Document structures vary, including unstructured free-text descriptions, semi-structured tabular data (e.g., MedDRA-coded adverse events, lab results), and structured patient demographics and medication history. Fields include common adverse event terms (e.g., CRS, ICANS grades), onset time, duration, and severity. Unique fields include CAR-T product batch numbers, infusion dosages, targets, co-stimulatory molecules, cytokine levels (e.g., IL-6, CRP), and neurotoxicity scale scores (e.g., ICE score). Units involve dosage (e.g., 10^6 cells/kg), time (days, hours), and cytokine concentration (pg/mL).
Constraints on Multi-Turn Conversations and Prompts
The high update frequency of CAR-T cell therapy data requires multi-turn conversation systems to quickly integrate the latest information. This prevents inaccurate responses based on outdated data. Diverse document structures, especially a large volume of unstructured text, mean prompt design must balance information extraction and semantic understanding. This ensures accurate identification of key adverse events and related details from free text. The presence of unique fields, such as CAR-T product batches and cytokine levels, means prompts must explicitly guide users to provide this information. The system must also link this information to standard adverse event terms for analysis. The use of specific domain terminology (e.g., CRS, ICANS) and scale scores (e.g., ICE score) requires the model to possess specialized medical knowledge. This ensures correct interpretation of professional vocabulary and accurate understanding and tracking of these concepts in multi-turn conversations.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
maxContext | 8 | Retains sufficient context in multi-turn conversations to cover critical patient treatment information and adverse event progression. |
temperature | 0.3 | Reduces randomness in model-generated responses, ensuring stable, objective, and fact-based answers in pharmacovigilance scenarios, avoiding hallucinations. |
Chunk size (Segment Length) | 500 characters (500 characters) | Accommodates the common medium-length descriptions found in clinical and individual case reports, ensuring semantic completeness and preventing truncation of key information. |
Recall count (Recall Count) | 10 | For complex queries, increasing the recall count improves the likelihood of retrieving relevant adverse event reports from vast pharmacovigilance data. |
Similarity threshold (Similarity Threshold) | 0.75 | Descriptions of CAR-T therapy adverse events can have subtle differences. A higher threshold helps filter out irrelevant recall results, improving precision. |
Rerank result count (Reranked Return Count) | 5 | After recalling multiple items, reranking places the most relevant few items at the forefront, improving response efficiency and accuracy. |
Common Pitfalls
- The model fails to accurately identify specific adverse event grades mentioned by the user, such as "Grade 3 CRS," leading to generic responses. This occurs when prompts do not explicitly guide the model to focus on and parse the severity and specific type of adverse event.
- Uploaded clinical trial report files cannot be effectively referenced or analyzed in the conversation, showing "file content empty." This happens when the
PARSE_FILE_TIMEOUT_SECONDSparameter in the file processing workflow is set too low, causing timeouts for large or complex documents. - When asked about adverse reactions for a specific CAR-T product batch, the model cannot link to historical data for that batch, responding "no relevant information found." This occurs when key identifiers like batch numbers are not effectively extracted and stored as queryable metadata during knowledge base construction.
Validation Steps
- Construct a set of test cases with multi-turn follow-up questions for typical CAR-T cell therapy adverse event scenarios. Observe if the model accurately identifies and correlates contextual information.
- Upload real adverse event reports containing complex free-text descriptions. Test if the model can accurately extract key adverse events, onset times, and severity information.
- Simulate user queries for adverse event data of a specific CAR-T product batch via the API. Verify if the returned results include adverse event records unique to that batch.
Note: The values provided are common starting points. Measure against your own samples to determine optimal settings.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.