Data Characteristics
CAR-T cell therapy product data originates from clinical trial reports, drug monographs, regulatory approval documents, academic papers, and patent literature. This data updates infrequently, typically with clinical trial progress, new drug approvals, or version revisions, with cycles ranging from months to years. Document structures are complex, containing extensive unstructured text, tables, charts, and biological sequence information. Fields and units are highly specialized; for example, dosage units may include cells/kg or IU/mL, and efficacy indicators like ORR (Objective Response Rate) and PFS (Progression-Free Survival) are common, often accompanied by specific statistical P-values and confidence intervals. Data also commonly includes macromolecular biological information such as gene sequences and protein structures.
Constraints Imposed by These Characteristics on Vector Models and Indexing
The highly specialized and complex nature of CAR-T cell therapy data demands more from vector models. Due to the large number of biological proper nouns and abbreviations, general text embedding models may struggle to accurately capture semantic relationships. Domain-specific customization or enhanced training may be necessary. The diversity of document structures, especially the extraction of information from charts and tables, increases the difficulty of text preprocessing, directly impacting segmentation and indexing quality. The specialized nature of fields and units requires vector models to distinguish similarities between different biological entities, for example, differentiating CAR-T products with different targets. Infrequent updates mean that once indexed, stability is high, but when new data arrives, efficient and accurate incremental updates are crucial. Special data types like gene sequences may require specific bioinformatics algorithms for preprocessing before vectorization.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Segment Length) | 512–768 characters (characters) | Balances semantic completeness with vector model input limits, preventing critical information dilution from overly long texts. |
Overlap Length | 128 characters (characters) | Ensures contextual continuity between paragraphs, preventing important information from being cut off. |
Vector Model (Vector Model) | text-embedding-ada-002 or domain-specific model | Balances generality with specialized domain performance, demonstrating good understanding of biomedical terminology. |
Similarity threshold (Similarity Threshold) | 0.75–0.85 | Balances recall rate and accuracy, avoiding the retrieval of irrelevant biomedical concepts. |
Recall count (Recall Count) | Top 8–15 entries (top 8–15 items) | Provides sufficient contextual information for subsequent language model processing while controlling query latency. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds (seconds) | Accommodates the parsing time for large clinical trial reports and patent documents, preventing parsing failures due to timeouts. |
Common Mistakes
- Knowledge base query results show low semantic relevance but high keyword matching. This typically occurs when using general vector models that fail to effectively understand the deep relationships between specialized terms and biological entities unique to the CAR-T domain.
- Uploading large PDF or Word format clinical reports results in a long system unresponsiveness or a
File Parsing Timeout(file parsing timeout) error. This happens when parameters likePARSE_FILE_TIMEOUT_SECONDSare not adjusted for the complex structure and extensive content of biomedical documents. - Despite correctly configuring
text-embedding-ada-002in the platform settings, ano available embedding modelerror still appears during knowledge base testing. This may be due to insufficient API key permissions or network connectivity issues, preventing the model service from being properly invoked.
Verification Steps
- Upload a drug monograph containing key information such as CAR-T dosage, targets, and adverse reactions. Then, query this information and observe whether the retrieved results contain correct and relevant segments, checking if their
similarityscores fall within the expected range. - Select multiple CAR-T products with similar biological mechanisms but different targets. Query each separately to verify if the vector model can distinguish these subtle differences, i.e., if the key characteristics of different products can be independently recalled.
- Gradually increase the number of documents in the knowledge base and record query response times. Ensure that query speed still meets real-time interaction requirements as data volume grows, and observe if
Recall count(recall count) remains stable. - Regularly check system logs to confirm that no
parsing failedorvectorization errormessages appear during file upload and indexing, especially for newly added complex documents.
Note: The values provided are common starting points. Measure against your own samples for optimal performance.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.