Multiturn Conversation and Prompts for CAR-T Cell Therapy Clinical Trial Pre-screening

CAR-T cell therapy clinical trial data originates from official databases like the China National Medical Products Administration (NMPA) Center for

Data Characteristics

CAR-T cell therapy clinical trial data originates from official databases like the China National Medical Products Administration (NMPA) Center for Drug Evaluation (CDE) clinical trial registration and information disclosure platform, and the U.S. National Institutes of Health (NIH) ClinicalTrials.gov. It also comes from medical journals, conference papers, and public pharmaceutical company reports. Data update frequencies vary. Official platforms typically update quarterly or monthly, while academic papers are continuously published. Document structures are primarily structured tabular data. They include patient inclusion criteria, exclusion criteria, treatment regimens, dosages, efficacy indicators (e.g., complete response rate CR, partial response rate PR), and adverse events AE. Unstructured text, such as study protocols, informed consent forms, and case report forms CRF, is also abundant. Fields and units strictly follow medical norms. For example, age is in "years," dosage is in "10^6 cells/kg" or "mg/kg," and adverse event grading uses CTCAE standards.

Constraints on Multiturn Conversation and Prompts

The specialized and multimodal nature of CAR-T clinical trial data significantly constrains multiturn conversation and prompt design. Precise queries for structured data require prompts to accurately map to database fields. An example is querying the "complete response rate" for patients with "CD19 positive B acute lymphoblastic leukemia." Unstructured text requires prompts with stronger semantic understanding to extract key inclusion and exclusion criteria from lengthy study protocols. Data update frequency determines the knowledge base maintenance cycle, ensuring conversations are based on the latest information. Medical terminology strictness requires prompts to avoid ambiguity. For example, "remission" can refer to complete remission, partial remission, or stable disease. This needs clarification in the conversation. Multiturn conversations need to track patient characteristics and treatment intent in context, progressively filtering for eligible clinical trials. An example is determining disease type in the first turn, then refining to specific genetic mutations or prior treatment history in the second turn.

Configuration Settings

Configuration ItemSuggested ValueRationale
maxContext8000 tokensAccommodates the length of medical text, ensuring complete context and preventing loss of critical information.
Chunk size (Chunk Length)500 characters (characters)Balances chunk granularity with semantic completeness, facilitating model understanding and retrieval.
Recall count (Recall Count)Top 8 entries (top 8)Increases coverage of relevant information to address the complexity of medical queries.
Similarity threshold (Similarity Threshold)0.78Ensures recalled results are highly relevant to the query, filtering out inaccurate or noisy information.
Rerank result count (Rerank Return Count)Top 5 entries (top 5)Selects the most relevant results for the user, improving information accuracy.
SYSTEM_PROMPTCalibrate based on actual measurementsMust include expert role setting for CAR-T clinical trial pre-screening and output format requirements.

Common Pitfalls

  • Irrelevant medical terms or data appear in the conversation. This occurs when prompts do not explicitly limit the knowledge base scope or lack strict intent recognition.
  • Model responses fail to accurately cite clinical trial inclusion/exclusion criteria or efficacy data. This occurs due to improper knowledge base chunking strategies, leading to truncated key information or semantic loss.
  • Results remain imprecise after multiturn conversations. An example is filtering for trials that do not meet specific adverse event AE grade requirements. This occurs when prompts have insufficient ability to parse complex logical conditions, failing to effectively translate user needs into queries.

Validation Steps

  • Execute multiturn conversations for typical patient profiles. Verify whether the final filtered clinical trial list fully complies with all inclusion and exclusion criteria.
  • Check if all data points cited by the model in the conversation (e.g., CR rate, AE grade, dosage) are consistent with the original knowledge base content, without fabrication or alteration.
  • Test queries with varying complexity of medical terminology. Verify whether the model can accurately identify and associate them with corresponding fields or text segments in the knowledge base.

Note: The values provided are common starting points and should be measured against your own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.