Deployment and Upgrade for CAR-T Cell Therapy Clinical Trial Pre-screening

CAR-T cell therapy clinical trial data primarily originates from global clinical trial registries (e.g., ClinicalTrials.gov, EU Clinical Trials

Data Characteristics

CAR-T cell therapy clinical trial data primarily originates from global clinical trial registries (e.g., ClinicalTrials.gov, EU Clinical Trials Register) and relevant academic journals and conference reports. Data update frequencies vary; registry information typically updates after trials reach specific stages, and some research data might lag by 6–12 months. Document structures are predominantly semi-structured and unstructured, including trial protocols, patient inclusion and exclusion criteria, treatment procedures, adverse event reports, and efficacy evaluation results. Fields include patient demographic information, disease diagnosis, CAR-T product type, dosage, pre-conditioning regimen, follow-up time, and primary and secondary endpoints (e.g., Objective Response Rate, Complete Response Rate, Progression-Free Survival, Overall Survival). Units are diverse; for example, dosage is expressed in 10^6 cells/kg, and follow-up time is in months or years.

Constraints Imposed by Data Characteristics on Deployment and Upgrade

The unstructured nature of CAR-T cell therapy data necessitates configuring robust text parsing capabilities during deployment to effectively extract key information. Documents are generally lengthy and contain extensive medical terminology, demanding higher requirements for segment length and number of recalled chunks to prevent loss of important context. Varying update frequencies mean the knowledge base must support incremental updates and version management to prevent outdated data from interfering with new information. The diversity of fields and inconsistent units require more refined adjustment of similarity threshold to balance recall and accuracy. For instance, pre-screening for different CAR-T product types requires accurate identification of product names and their associated clinical indicators. When processing this data, the model needs strong semantic understanding capabilities to handle the complexity and polysemy of medical terminology.

Configuration Settings

Configuration ItemSuggested ValueRationale
UPLOAD_FILE_MAX_SIZE500 MBClinical trial protocols and reports can be large; support for uploading large files is needed.
PARSE_FILE_TIMEOUT_SECONDS600Parsing large PDF files can be time-consuming; sufficient time must be allocated.
segment length800–1200 charactersMedical documents have strong contextual relevance; longer segments retain semantic meaning.
number of recalled chunkstop 8–12 chunksEnsures coverage of multi-source information, improving key information capture rate.
similarity threshold0.75–0.85Balances precise matching of medical terms with identification of potential variations.
number of reranked resultstop 5Prioritizes the most relevant core information, improving pre-screening efficiency.

Common Mistakes

  • After an upgrade, question splitting performance degrades, leading to inaccurate pre-screening results. This often occurs because the new version's tokenizer or text processing logic differs from the old one, requiring re-evaluation and adjustment of parameters like segment length.
  • A locally deployed model errors out with a 403 status code after adding a knowledge base. This could be due to incorrect token permission configuration, or the knowledge base file path or format not meeting model requirements, preventing the model from accessing or parsing the knowledge base data.
  • PDF file parsing fails, and logs show an error message. This is typically related to PARSE_FILE_TIMEOUT_SECONDS being set too short, or the PDF file itself containing complex charts or scanned images, requiring a more powerful parsing engine.

How to Verify Correct Configuration

  • Upload a typical CAR-T clinical trial protocol PDF file and check if the file status in the knowledge base documents list is "parsed successfully".
  • Use a query containing a specific CAR-T product name and treatment indicators to verify that the recall results include relevant segments from multiple documents, and check if the number of recalled chunks meets expectations.
  • Enter a medical term with a typo or synonym and observe if the similarity threshold effectively identifies and recalls relevant information.
  • For complex queries such as patient inclusion criteria, check if the content in the number of reranked results is highly focused on key screening conditions.

The values provided are common starting points and should be measured against the reader's own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.