Data Characteristics
CAR-T cell therapy clinical trial data originates primarily from clinical trial registries (e.g., ClinicalTrials.gov, Chinese Clinical Trial Registry) and academic publications, conference abstracts, and internal research reports from various institutions. This data updates frequently, especially with new targets and therapies. Document structures typically include trial protocols, patient inclusion/exclusion criteria, treatment regimens, adverse event reports (AEs), efficacy endpoints (e.g., ORR, CR rate), and follow-up data. Fields cover medical terminology (e.g., ICD-O-3 tumor codes), drug dosages (e.g., cells/kg), time units (e.g., months, weeks), biomarkers (e.g., CD19 expression levels), and gene mutation information.
Constraints from "Vector Models and Indexing"
CAR-T clinical trial data is highly specialized, containing extensive medical jargon, abbreviations, and numerical ranges. This requires vector models to have strong semantic understanding capabilities. Patient inclusion/exclusion criteria often involve complex logical relationships, such as "patients who have previously received specific treatments are excluded, except for lymphoma relapse patients." This demands that vector indexing captures subtle differences between statements. Treatment regimens and adverse event reports are usually unstructured text, while dosage and efficacy data are often structured or semi-structured. Frequent updates necessitate a fast incremental update mechanism for the index to avoid frequent full rebuilds. Patient privacy and data security are core considerations; the indexing process must ensure sensitive information is de-identified.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Chunk Size) | 500–800 characters | Balances semantic completeness with vector model processing efficiency. Avoids overly long chunks diluting key information and overly short chunks losing context. |
Chunk Overlap Length (Overlap Length) | 100–150 characters | Ensures contextual continuity across chunks, especially when describing inclusion/exclusion criteria or treatment procedures. |
Recall count (Recall Count) | Top 10–15 items | Given the complexity of CAR-T trials, more relevant document segments are needed for reranking to ensure comprehensive coverage. |
Similarity threshold (Similarity Threshold) | Calibrate based on actual measurements | Fine-tune between 0.75 and 0.85 through A/B testing, based on specific query scenarios and dataset characteristics, to balance recall and precision. |
Index Model (Embedding Model) | text-embedding-3-large or equivalent | Addresses medical terminology and complex logic, requiring a high-performance embedding model to capture fine-grained semantics. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Allows sufficient time for parsing large clinical trial protocol PDF documents, which can be time-consuming. |
Common Pitfalls
- The index status remains "indexing" for an extended period, failing to transition to "ready." Common reasons include document parsing timeouts or slow vector model service responses.
- The number of data entries or indexes in the dataset increases abnormally. This usually results from duplicate uploads of the same document or improperly configured scheduled synchronization tasks leading to data redundancy.
- Index building tasks stall after switching to a high-performance embedding model. This may be due to insufficient model service resources or API call frequency limits.
Verification Steps
- Upload a clinical trial protocol PDF file containing complex inclusion/exclusion criteria. Check if the index builds successfully and verify recall results by retrieving relevant keywords.
- Use query statements containing specific dosage units (e.g.,
mg/kg) and time units (e.g.,weeks). Verify the system accurately recalls relevant document segments. - Simulate high-frequency data update scenarios. Observe if incremental indexing tasks complete promptly and check if new data is retrievable.
Note: The values provided are common starting points. Measure them against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.