Knowledge Base Retrieval and Recall for CAR-T Cell Therapy Clinical Trial Pre-screening

CAR-T cell therapy clinical trial data originates primarily from clinical trial registration platforms (e.g., ClinicalTrials.gov, Chinese Clinical

Data Characteristics

CAR-T cell therapy clinical trial data originates primarily from clinical trial registration platforms (e.g., ClinicalTrials.gov, Chinese Clinical Trial Registry), investigator brochures, medical journal papers, and pharmaceutical company internal reports. Data updates frequently, typically weekly or monthly, triggered by new trial registrations, changes in patient recruitment status, and research result publications. Document structures are predominantly semi-structured and unstructured. Semi-structured data includes trial protocols, patient inclusion/exclusion criteria, and adverse event reports, containing explicit fields such as NCT Number, Indication, Intervention, Disease Stage, and Prior Treatment History. Unstructured data exists in trial descriptions, research backgrounds, and discussion sections. Field and unit specificities include dosage units often expressed as cells/kg or cells/m², and time units detailed to weeks, months, years, frequently accompanied by concepts like follow-up period.

Constraints on Knowledge Base Retrieval and Recall

The high update frequency of CAR-T clinical trial data requires the knowledge base to have efficient incremental update mechanisms to ensure retrieval result timeliness. The presence of semi-structured data means traditional full-text search is insufficient for precise matching, necessitating integration with structured information for filtering and ranking. For example, precisely retrieving trials for a specific Disease Stage is difficult using only text similarity. Extracting key information from unstructured data, such as complex descriptions of Prior Treatment History, demands advanced tokenization and entity recognition. Detailed dosage and time units require support for numerical range queries and unit conversions, such as retrieving trials with dosage greater than 5x10^6 cells/kg. Furthermore, multiple conditional combinations within patient inclusion/exclusion criteria require retrieval logic capable of handling complex Boolean operations and nested conditions to precisely match pre-screening requirements.

Configuration Guidelines

Configuration ItemRecommended ValueRationale
Chunk size (Segment Length)500–800 charactersClinical trial document paragraphs often contain multiple key pieces of information; segments that are too short risk losing context, while segments that are too long introduce noise.
Recall count (Recall Count)10–15 itemsInitial recall needs to cover a sufficient number of potentially relevant trials for subsequent re-ranking and filtering.
Similarity threshold (Similarity Threshold)Calibrated by actual measurementFor CAR-T trial specific terminology and abbreviations, the threshold needs to balance recall and precision.
Rerank result count (Re-ranked Return Count)3–5 itemsThe final results presented to the user should be highly relevant and concise, avoiding information overload.
Max File Size100 MBConsiders the volume of investigator brochures and detailed reports, ensuring large PDF documents can be uploaded.
Embedding Modelbge-large-zh or text-embedding-ada-002Selects an embedding model that performs well in the medical domain to improve semantic understanding of specialized terminology.

Common Pitfalls

  • Retrieval results contain many irrelevant trials. This occurs when Recall count (Recall Count) is high but relevance is low. The Similarity threshold (Similarity Threshold) might be set too low, leading to the recall of many low-relevance documents, or the segmentation strategy might be inappropriate, causing key information to be fragmented.
  • When querying for trials with specific dosages, the system fails to correctly identify numerical ranges or units, resulting in empty results. This happens if the knowledge base lacks specialized entity recognition or parsing configurations for CAR-T-specific dosage units and numerical ranges.
  • Updated clinical trial data is not reflected in retrieval results in a timely manner. This manifests as latest status not matching reality. The knowledge base's incremental synchronization mechanism might not be effectively configured or triggered, leading to outdated data.

Verification Steps

  • Select a batch of CAR-T clinical trial cases with clear inclusion/exclusion criteria. Construct query statements and check if the trials in the Rerank result count (Re-ranked Return Count) accurately match.
  • For queries involving different dosage ranges and units, such as CD19 CAR-T dosage > 5e6 cells/kg, verify if the system correctly recalls target trials and check the parsing accuracy of the dosage field.
  • Simulate a CAR-T clinical trial status update (e.g., recruitment status changing from Recruiting to Completed). Query immediately after the knowledge base update to confirm if the status change has taken effect.

The values given are common starting points and should be measured against the reader's own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.