CAR-T Cell Therapy Clinical Trial Pre-screening: Citation and Traceability

CAR-T cell therapy clinical trial data originates from global clinical trial registries (e.g., ClinicalTrials.gov, European Clinical Trials Register)

Data Characteristics

CAR-T cell therapy clinical trial data originates from global clinical trial registries (e.g., ClinicalTrials.gov, European Clinical Trials Register) and academic journals or conference reports. This data updates frequently, typically quarterly or semi-annually, as trials progress or results publish. Document structures are a mix of structured tables and unstructured text. Structured sections include trial ID, study phase, disease type, drug name, dosage, patient inclusion/exclusion criteria, and primary/secondary endpoints. Unstructured sections cover detailed trial protocols, patient recruitment information, adverse event reports, and researcher interpretations of results. Fields include gene editing sites, cell expansion folds, infusion dosage (e.g., 1x10^6 CAR-T cells/kg), and follow-up duration (e.g., 6 months, 12 months). Units for cell count, dosage, time, and percentage require strict differentiation.

Constraints for Citation and Traceability

High update frequency of CAR-T clinical trial data requires the knowledge base to quickly synchronize with the latest developments, preventing the citation of outdated information. The mixed structured and unstructured document format necessitates extracting precise numerical values from structured fields and understanding complex inclusion/exclusion criteria from unstructured text during pre-screening. For example, critical information like patient genotype, prior treatment history, and specific biomarker expression levels may be scattered across multiple descriptive paragraphs. Citations must precisely pinpoint these specific paragraphs or table rows to support decisions on patient eligibility. The highly specialized nature of CAR-T therapy makes source authority crucial. The system must trace back to the original publishing institution or research team to ensure information reliability and identify potential trial design differences or reporting biases.

Configuration Settings

Configuration ItemRecommended ValueRationale
Chunk size (Chunk Size)500–800 charactersAccommodates detailed inclusion/exclusion criteria and adverse event descriptions in CAR-T clinical trial protocols, ensuring contextual completeness.
Recall count (Recall Count)Top 8Considers the high complexity of CAR-T trials, requiring more relevant paragraphs for comprehensive judgment.
Similarity threshold (Similarity Threshold)0.75Increases matching precision, avoiding confusion between trials with different targets or treatment stages.
Rerank result count (Rerank Return Count)3Further refines recalled results, focusing on the three most relevant citations to reduce redundant information.
Citation Variable Format[Title](URL)Directly provides original document links, allowing engineers to quickly access full trial reports.
Show CitationsDisplayCritical information requires clear sources to support the rigor of clinical judgments.

Common Pitfalls

  • Missing citations in model responses may stem from an improper knowledge base chunking strategy, leading to truncated key information and ineffective citations.
  • The system prompt quote type error typically indicates a mismatch between the variable type passed and the expected type, such as passing a text description when a numerical value is expected.
  • Quoted trial versions in responses not aligning with the latest developments indicate that the knowledge base synchronization mechanism failed to update in time, leading to citations of outdated clinical trial data.

Validation Steps

  • Input multiple typical patient cases and verify that the cited trial ID, disease type, and inclusion/exclusion criteria in the response are accurate.
  • Check if the citation source URLs are accessible and point to the original clinical trial registration page or academic literature.
  • Compare key numerical values extracted in the model response (e.g., CD19 positivity rate, dosage range) with the original quoted text to confirm consistency.

The values provided are common starting points. Measure them against your own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.