Data Characteristics
R&D document data in CAR-T cell therapy is unique. Data sources include clinical trial reports, research papers, patent literature, bioinformatics databases (e.g., GenBank, UniProt), and internal experimental records. These documents update relatively quickly, especially clinical trials and research progress, with potential quarterly or monthly updates. Document structures often contain highly specialized biomedical terminology, gene sequences, protein structures, cell pathway diagrams, drug mechanism descriptions, dose-response curves, and more. Fields and units are complex, such as gene names (CD19, BCMA), protein IDs, cell concentrations (cells/µL), dosages (mg/kg), time points (days, weeks, months), response rates (%), and various statistical indicators.
Constraints from Data Characteristics on Vector Models and Indexing
The specialized nature and complex structure of CAR-T cell therapy documents impose specific requirements on vector models and indexing. Extensive specialized terminology and abbreviations require glossaries and ontologies to ensure accurate semantic capture during vectorization. Non-textual information like gene sequences and protein structures need specialized encoding or embedding methods, or conversion into descriptive text. High document update frequency means the index needs to support efficient incremental update mechanisms to avoid frequent full rebuilds. Diverse fields and units require vector models to distinguish between numerical and descriptive information and support unit and range-based filtering during retrieval. Additionally, common charts and image information in documents, while currently relying on text descriptions, should consider their potential structured information.
Configuration Settings
| Configuration Item | Suggested Value | Rationale |
|---|---|---|
Chunk size (Segment Length) | 500-800 characters (characters) | CAR-T documents contain extensive, logically dense professional descriptions. Segments that are too short may cut off critical information, while segments that are too long may introduce excessive noise, affecting the precision of vector representation. |
Chunk Overlap Length (Segment Overlap Length) | 100-150 characters (characters) | Ensures that critical concepts and context are not lost across segments, especially for gene pathways and clinical outcome descriptions. Overlapping parts help maintain semantic coherence. |
Vector Model (Vector Model) | text-embedding-ada-002 or domain-specific model | Given the specialized nature of the biomedical field, choose a foundational model trained on a large volume of text, or a model fine-tuned on biomedical corpora, to better understand specialized terminology and conceptual relationships. |
Similarity threshold (Similarity Threshold) | 0.75-0.85 | The CAR-T field demands high retrieval accuracy. A threshold that is too low may introduce irrelevant results, while one that is too high may miss potentially related information. Specific values require actual testing, calibrated against recall and precision. |
Recall count (Recall Count) | Top 10-15 entries (top 10-15 items) | The initial recall stage needs to cover as many potentially relevant document segments as possible for subsequent re-ranking model selection. Considering the complexity of document content, appropriately increasing the recall count can reduce the risk of omissions. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds (seconds) | R&D documents, especially clinical trial reports and patents, can be large, making parsing time-consuming. Set a sufficiently long timeout to prevent indexing failures due to file parsing interruptions. |
Common Mistakes
- Knowledge base index is not ready, leading to empty or irrelevant retrieval results. This happens when document parsing fails or the vector generation process is abnormal, preventing correct conversion of document content into retrievable vectors.
- Retrieval results contain a large amount of irrelevant gene or protein information. This typically results from an unreasonable segmentation strategy, which fails to effectively isolate different entity descriptions, or the vector model has insufficient discriminative power for specific biological terms.
- System error "context length exceeded" or incomplete context information is returned. This often occurs when the segment length is set too large, causing a single segment to exceed the processing limit of the vector model or subsequent language model, or when document context is not fully considered during retrieval.
Verification Steps
- Upload a batch of representative CAR-T R&D documents. Observe parsing logs to confirm successful file parsing status.
- Perform a series of queries on the knowledge base, including specialized terms and complex concepts. Check if the "Recall Count" of the returned results meets expectations, and manually evaluate the relevance of the top few results.
- Query for specific genes, targets, or clinical trial codes. Verify that the returned document snippets accurately contain these entities and their contextual information, and assess if the similarity scores are within a reasonable range.
- Regularly check the completion time and resource consumption of index building tasks to ensure stable operation of the incremental update mechanism.
The values given are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.