Multi-Turn Conversations and Prompts for CAR-T Cell Therapy Quality Documentation

CAR-T cell therapy quality documentation covers the entire lifecycle, from cell collection, genetic modification, ex vivo expansion, and formulation

Data Characteristics of This Category

CAR-T cell therapy quality documentation covers the entire lifecycle, from cell collection, genetic modification, ex vivo expansion, and formulation manufacturing to quality inspection and release. These documents typically exist in PDF, Word, and Excel formats, including Standard Operating Procedures (SOPs), Batch Production Records (BPRs), Batch Quality Control Records (BJRs), equipment calibration records, deviation handling, change control, and risk assessment reports. Data sources are extensive, involving multiple departments and instruments. Document update frequency is high, especially for SOPs and batch records, which are revised with process optimization, regulatory updates, or deviation handling. Document structure is complex, containing numerous tables, charts, flowcharts, and specialized terminology. Fields and units are highly specific, such as cell count (×10^6 cells/mL), viral titer (TU/mL), transduction efficiency (%), purity (%), and viability (%), requiring strict adherence to numerical precision and unit consistency.

Constraints Imposed by These Characteristics on "Multi-Turn Conversations and Prompts"

The data characteristics of CAR-T cell therapy quality documentation impose specific constraints on multi-turn conversation and prompt design. The highly specialized content and the presence of complex tables and charts require the RAG model to effectively parse non-textual information and maintain accurate understanding of specialized terminology in multi-turn conversations, avoiding semantic drift. Frequent document updates mean the knowledge base needs to support efficient incremental updates and version management, ensuring conversations are based on the latest data. The large volume of numerical data and strict unit requirements in batch and inspection records necessitate prompt designs that guide the model to focus on numerical accuracy and unit matching. For example, when querying a specific quality control metric for a particular batch, the model must be able to extract and display it precisely. Furthermore, multi-turn conversations exhibit a stronger dependency on historical context. For instance, a user might first ask about an SOP and then inquire about the execution records of a specific step within that SOP, requiring the system to have stronger contextual association capabilities.

Configuration Settings

Configuration ItemSuggested ValueRationale
Chunk size (Chunk Size)800–1200 charactersBalances the dense specialized terminology of CAR-T documents with logical coherence between paragraphs, avoiding semantic integrity loss during splitting.
Recall count (Recall Count)Top 8–12 entriesThe complexity of CAR-T documents means a single query may involve multiple relevant paragraphs; increasing recall improves coverage.
Similarity threshold (Similarity Threshold)0.78–0.85Ensures recalled document chunks are highly relevant to the query intent, filtering out significant medical background noise.
Rerank result count (Reranked Return Count)Top 5 entriesAfter high recall, reranking further improves the ranking of the most relevant information, optimizing conversation quality.
maxContext6000–8000 tokensAccommodates the accumulation of specialized terminology and contextual information in multi-turn conversations, preventing premature truncation.
embeddingModeltext-embedding-ada-002 or higher versionHandles specialized vocabulary and complex concepts in the CAR-T domain, improving vectorization accuracy.

Three Common Pitfalls

  • Misunderstanding or confusion of numerous specialized terms in conversations, such as confusing "transduction efficiency" with "expansion fold." This occurs due to excessively small chunk sizes or insufficient understanding of domain-specific vocabulary by the embedding model.
  • When a user asks for specific test results for a particular batch, the model returns data from other batches or misses critical numerical values. This happens because document parsing fails to accurately extract row and column association information from tables.
  • After prolonged multi-turn conversations, the model starts "forgetting" earlier conversation content, leading to off-topic responses or repeated questions. This is due to maxContext being set too low, unable to accommodate the complete conversation history.

How to Confirm Proper Configuration

  • Build test sets for different types of CAR-T quality documents (SOPs, batch records, quality inspection reports). Test the model's accuracy in extracting key information (e.g., specific quality control parameter values, detailed operating steps) and compare with manual verification results.
  • Simulate multi-turn conversation scenarios. For example, first ask for the name of an SOP, then follow up with the calibration frequency of a specific piece of equipment within that SOP. Check if the model can accurately associate context and provide coherent and correct answers.
  • Randomly select documents containing tables and charts. Test if the model can correctly parse and cite data from them in conversations. For instance, inquire about the viral titer of a specific batch and check if the returned numerical value and unit match the original text.

The values provided are common starting points and should be measured against the reader's own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.