Knowledge Base Retrieval and Recall for CAR-T Cell Therapy Products

CAR-T cell therapy data originates from clinical trial reports, drug labels, regulatory approval documents, academic papers, and patent literature.

Data Characteristics

CAR-T cell therapy data originates from clinical trial reports, drug labels, regulatory approval documents, academic papers, and patent literature. This data updates frequently, especially clinical trial progress and new drug approval information. Document structures typically include detailed drug mechanisms of action, indications, dosage and administration, adverse reactions, manufacturing processes, and quality control standards. Fields often involve gene sequences, protein structures, cell counts, treatment dosages, treatment courses, patient inclusion and exclusion criteria, clinical endpoints (e.g., complete response rate CR, overall survival OS), and adverse event classifications (e.g., CRS, ICANS). Units include international units IU, milligrams mg, milliliters mL, cell counts cells, percentages %, and time units (days, months, years).

Constraints on Knowledge Base Retrieval and Recall

The specialized and complex nature of CAR-T cell therapy data requires precise identification of biomedical terms and abbreviations during knowledge base retrieval. For example, distinguishing whether CR refers to complete response rate or another meaning. The rapid update frequency necessitates efficient incremental update mechanisms to ensure retrieval results are current. Documents containing charts, tables, and structured data challenge text extraction and semantic understanding; plain text retrieval may not capture all key information. Furthermore, precise numerical values like gene sequences and cell counts require retrieval to match text and filter by numerical ranges or specific thresholds. Standardized adverse event classifications require the retrieval system to understand and associate similar risks under different terminologies.

Configuration Guidelines

Configuration ItemSuggested ValueRationale
Chunk Length800–1200 charactersEnsures key information (e.g., mechanism, adverse reactions) is within one chunk
Recall CountTop 15Covers more potentially relevant documents, addresses terminology diversity
Similarity Threshold0.75Balances retrieval precision and recall rate, avoids over-generalization
Rerank Return CountTop 5Selects the most relevant results, improves user experience
UPLOAD_FILE_MAX_SIZE500 MBAccommodates large clinical trial reports or multimedia attachments
PARSE_FILE_TIMEOUT_SECONDS600 secondsEnsures complex PDFs or scanned documents have sufficient time to parse

Common Pitfalls

  • During knowledge base training, a "partial file processing error" may occur. This can happen if PDF documents contain numerous scanned images, leading to OCR recognition failure or timeout.
  • Retrieval results may contain many irrelevant entries. This indicates the Similarity Threshold is too low, failing to effectively filter out noise.
  • Knowledge base search response times may be excessively long. Logs showing abnormal query_latency_ms suggest vector database index parameters (e.g., hnsw.max_scan_tuple) are not optimized, leading to inefficient queries.

Validation Steps

  • Upload a batch of CAR-T related documents with various file types (PDF, DOCX, TXT). Check if all files are successfully trained and ingested into the knowledge base by observing the File Status field.
  • Use a series of queries containing specialized terms, abbreviations, and numerical values. Check the relevance of retrieval results and compare Similarity Scores.
  • Simulate high-concurrency query scenarios. Monitor system Response Time and CPU usage to ensure performance meets expectations.

The values provided are common starting points and should be measured against the reader's own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.