Data Characteristics
CAR-T cell therapy clinical trial data primarily originates from global clinical trial registries (e.g., ClinicalTrials.gov, EU Clinical Trials Register) and relevant academic journals and conference reports. Data update frequencies vary; registry information typically updates after trials reach specific stages, and some research data might lag by 6–12 months. Document structures are predominantly semi-structured and unstructured, including trial protocols, patient inclusion and exclusion criteria, treatment procedures, adverse event reports, and efficacy evaluation results. Fields include patient demographic information, disease diagnosis, CAR-T product type, dosage, pre-conditioning regimen, follow-up time, and primary and secondary endpoints (e.g., Objective Response Rate, Complete Response Rate, Progression-Free Survival, Overall Survival). Units are diverse; for example, dosage is expressed in 10^6 cells/kg, and follow-up time is in months or years.
Constraints Imposed by Data Characteristics on Deployment and Upgrade
The unstructured nature of CAR-T cell therapy data necessitates configuring robust text parsing capabilities during deployment to effectively extract key information. Documents are generally lengthy and contain extensive medical terminology, demanding higher requirements for segment length and number of recalled chunks to prevent loss of important context. Varying update frequencies mean the knowledge base must support incremental updates and version management to prevent outdated data from interfering with new information. The diversity of fields and inconsistent units require more refined adjustment of similarity threshold to balance recall and accuracy. For instance, pre-screening for different CAR-T product types requires accurate identification of product names and their associated clinical indicators. When processing this data, the model needs strong semantic understanding capabilities to handle the complexity and polysemy of medical terminology.
Configuration Settings
| Configuration Item | Suggested Value | Rationale |
|---|---|---|
UPLOAD_FILE_MAX_SIZE | 500 MB | Clinical trial protocols and reports can be large; support for uploading large files is needed. |
PARSE_FILE_TIMEOUT_SECONDS | 600 | Parsing large PDF files can be time-consuming; sufficient time must be allocated. |
segment length | 800–1200 characters | Medical documents have strong contextual relevance; longer segments retain semantic meaning. |
number of recalled chunks | top 8–12 chunks | Ensures coverage of multi-source information, improving key information capture rate. |
similarity threshold | 0.75–0.85 | Balances precise matching of medical terms with identification of potential variations. |
number of reranked results | top 5 | Prioritizes the most relevant core information, improving pre-screening efficiency. |
Common Mistakes
- After an upgrade,
question splittingperformance degrades, leading to inaccurate pre-screening results. This often occurs because the new version's tokenizer or text processing logic differs from the old one, requiring re-evaluation and adjustment of parameters likesegment length. - A locally deployed model errors out with a
403status code after adding a knowledge base. This could be due to incorrecttokenpermission configuration, or the knowledge base file path or format not meeting model requirements, preventing the model from accessing or parsing the knowledge base data. - PDF file parsing fails, and logs show an
error message. This is typically related toPARSE_FILE_TIMEOUT_SECONDSbeing set too short, or the PDF file itself containing complex charts or scanned images, requiring a more powerful parsing engine.
How to Verify Correct Configuration
- Upload a typical CAR-T clinical trial protocol PDF file and check if the file status in the
knowledge base documentslist is "parsed successfully". - Use a query containing a specific CAR-T product name and treatment indicators to verify that the recall results include relevant segments from multiple documents, and check if the
number of recalled chunksmeets expectations. - Enter a medical term with a typo or synonym and observe if the
similarity thresholdeffectively identifies and recalls relevant information. - For complex queries such as patient inclusion criteria, check if the content in the
number of reranked resultsis highly focused on key screening conditions.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.