Data Characteristics
CAR-T cell therapy data originates from clinical trial reports, drug labels, regulatory approval documents, academic papers, and patent literature. This data updates frequently, especially clinical trial progress and new drug approval information. Document structures typically include detailed drug mechanisms of action, indications, dosage and administration, adverse reactions, manufacturing processes, and quality control standards. Fields often involve gene sequences, protein structures, cell counts, treatment dosages, treatment courses, patient inclusion and exclusion criteria, clinical endpoints (e.g., complete response rate CR, overall survival OS), and adverse event classifications (e.g., CRS, ICANS). Units include international units IU, milligrams mg, milliliters mL, cell counts cells, percentages %, and time units (days, months, years).
Constraints on Knowledge Base Retrieval and Recall
The specialized and complex nature of CAR-T cell therapy data requires precise identification of biomedical terms and abbreviations during knowledge base retrieval. For example, distinguishing whether CR refers to complete response rate or another meaning. The rapid update frequency necessitates efficient incremental update mechanisms to ensure retrieval results are current. Documents containing charts, tables, and structured data challenge text extraction and semantic understanding; plain text retrieval may not capture all key information. Furthermore, precise numerical values like gene sequences and cell counts require retrieval to match text and filter by numerical ranges or specific thresholds. Standardized adverse event classifications require the retrieval system to understand and associate similar risks under different terminologies.
Configuration Guidelines
| Configuration Item | Suggested Value | Rationale |
|---|---|---|
Chunk Length | 800–1200 characters | Ensures key information (e.g., mechanism, adverse reactions) is within one chunk |
Recall Count | Top 15 | Covers more potentially relevant documents, addresses terminology diversity |
Similarity Threshold | 0.75 | Balances retrieval precision and recall rate, avoids over-generalization |
Rerank Return Count | Top 5 | Selects the most relevant results, improves user experience |
UPLOAD_FILE_MAX_SIZE | 500 MB | Accommodates large clinical trial reports or multimedia attachments |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Ensures complex PDFs or scanned documents have sufficient time to parse |
Common Pitfalls
- During knowledge base training, a "partial file processing error" may occur. This can happen if PDF documents contain numerous scanned images, leading to OCR recognition failure or timeout.
- Retrieval results may contain many irrelevant entries. This indicates the
Similarity Thresholdis too low, failing to effectively filter out noise. - Knowledge base search response times may be excessively long. Logs showing abnormal
query_latency_mssuggest vector database index parameters (e.g.,hnsw.max_scan_tuple) are not optimized, leading to inefficient queries.
Validation Steps
- Upload a batch of CAR-T related documents with various file types (PDF, DOCX, TXT). Check if all files are successfully trained and ingested into the knowledge base by observing the
File Statusfield. - Use a series of queries containing specialized terms, abbreviations, and numerical values. Check the relevance of retrieval results and compare
Similarity Scores. - Simulate high-concurrency query scenarios. Monitor system
Response TimeandCPUusage to ensure performance meets expectations.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.