Data Characteristics
CAR-T cell therapy clinical trial data originates primarily from clinical trial registration platforms (e.g., ClinicalTrials.gov, Chinese Clinical Trial Registry), investigator brochures, medical journal papers, and pharmaceutical company internal reports. Data updates frequently, typically weekly or monthly, triggered by new trial registrations, changes in patient recruitment status, and research result publications. Document structures are predominantly semi-structured and unstructured. Semi-structured data includes trial protocols, patient inclusion/exclusion criteria, and adverse event reports, containing explicit fields such as NCT Number, Indication, Intervention, Disease Stage, and Prior Treatment History. Unstructured data exists in trial descriptions, research backgrounds, and discussion sections. Field and unit specificities include dosage units often expressed as cells/kg or cells/m², and time units detailed to weeks, months, years, frequently accompanied by concepts like follow-up period.
Constraints on Knowledge Base Retrieval and Recall
The high update frequency of CAR-T clinical trial data requires the knowledge base to have efficient incremental update mechanisms to ensure retrieval result timeliness. The presence of semi-structured data means traditional full-text search is insufficient for precise matching, necessitating integration with structured information for filtering and ranking. For example, precisely retrieving trials for a specific Disease Stage is difficult using only text similarity. Extracting key information from unstructured data, such as complex descriptions of Prior Treatment History, demands advanced tokenization and entity recognition. Detailed dosage and time units require support for numerical range queries and unit conversions, such as retrieving trials with dosage greater than 5x10^6 cells/kg. Furthermore, multiple conditional combinations within patient inclusion/exclusion criteria require retrieval logic capable of handling complex Boolean operations and nested conditions to precisely match pre-screening requirements.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Segment Length) | 500–800 characters | Clinical trial document paragraphs often contain multiple key pieces of information; segments that are too short risk losing context, while segments that are too long introduce noise. |
Recall count (Recall Count) | 10–15 items | Initial recall needs to cover a sufficient number of potentially relevant trials for subsequent re-ranking and filtering. |
Similarity threshold (Similarity Threshold) | Calibrated by actual measurement | For CAR-T trial specific terminology and abbreviations, the threshold needs to balance recall and precision. |
Rerank result count (Re-ranked Return Count) | 3–5 items | The final results presented to the user should be highly relevant and concise, avoiding information overload. |
Max File Size | 100 MB | Considers the volume of investigator brochures and detailed reports, ensuring large PDF documents can be uploaded. |
Embedding Model | bge-large-zh or text-embedding-ada-002 | Selects an embedding model that performs well in the medical domain to improve semantic understanding of specialized terminology. |
Common Pitfalls
- Retrieval results contain many irrelevant trials. This occurs when
Recall count(Recall Count) is high but relevance is low. TheSimilarity threshold(Similarity Threshold) might be set too low, leading to the recall of many low-relevance documents, or the segmentation strategy might be inappropriate, causing key information to be fragmented. - When querying for trials with specific dosages, the system fails to correctly identify numerical ranges or units, resulting in
empty results. This happens if the knowledge base lacks specialized entity recognition or parsing configurations for CAR-T-specific dosage units and numerical ranges. - Updated clinical trial data is not reflected in retrieval results in a timely manner. This manifests as
latest statusnot matching reality. The knowledge base's incremental synchronization mechanism might not be effectively configured or triggered, leading to outdated data.
Verification Steps
- Select a batch of CAR-T clinical trial cases with clear inclusion/exclusion criteria. Construct query statements and check if the trials in the
Rerank result count(Re-ranked Return Count) accurately match. - For queries involving different dosage ranges and units, such as
CD19 CAR-T dosage > 5e6 cells/kg, verify if the system correctly recalls target trials and check the parsing accuracy of thedosagefield. - Simulate a CAR-T clinical trial status update (e.g., recruitment status changing from
RecruitingtoCompleted). Query immediately after the knowledge base update to confirm if the status change has taken effect.
The values given are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.