Data Characteristics for This Category
CAR-T cell therapy is a highly specialized, novel treatment. Its regulations and Standard Operating Procedure (SOP) documents have distinct characteristics. Data sources primarily include regulations and guidelines from drug regulatory agencies, clinical trial protocols, hospital-developed operating procedures, and pharmaceutical company product inserts. These documents update infrequently, typically only with policy adjustments or technological advancements. Document structures are complex, often containing extensive specialized terminology, flowcharts, tables, and references. Fields include, but are not limited to: cell preparation standards, patient screening criteria, administration procedures, adverse event management, and follow-up requirements. Units involve dosage (e.g., cells/kg), time (e.g., hours, days), temperature (e.g., ℃), and require extremely high precision.
Constraints Imposed by These Characteristics on "Multi-Turn Conversations and Prompts"
The specialized nature, complexity, and high accuracy requirements of CAR-T cell therapy regulatory documents impose strict constraints on multi-turn conversation and prompt design. First, multi-turn conversations must accurately understand specialized terms in user queries to avoid incorrect answers due to semantic deviations. Second, the highly procedural information in SOPs requires the model to understand logical relationships between steps, supporting users in progressively deeper queries across multiple interactions. The low update frequency of documents means historical data is relatively stable, but once updated, the knowledge base must refresh synchronously to avoid providing outdated information. Precision requirements for fields and units necessitate prompt designs that guide the model to focus on numerical details and present them accurately in responses, for example, explicitly stating units when discussing dosage.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
maxContext | 2048 tokens | Balances long conversation history with model processing capability, preventing single-turn context overflow. |
temperature | 0.1 | Ensures rigor and accuracy of responses, reduces hallucination risk, aligning with regulatory Q&A characteristics. |
top_p | 0.3 | Further limits model generation diversity, focusing on factual content. |
Chunk size (Segment Length) | 500–700 characters | Accommodates the paragraph length of regulatory documents, ensuring semantic completeness and reducing context loss. |
Recall count (Recall Count) | Top 5 entries (Top 5) | Balances recall efficiency and relevance, covering core information points. |
Similarity threshold (Similarity Threshold) | 0.85 | Improves the precision of recall results, ensuring only highly relevant document snippets are retrieved. |
Three Common Pitfalls
- The model misunderstands specialized terms in multi-turn conversations, leading to inaccurate responses. This typically occurs due to incomplete term definitions in the knowledge base or prompts that insufficiently guide the model to focus on professional semantics.
- When users ask a second or subsequent question, the model fails to correctly associate with previous context, resulting in disjointed responses. This might be because the
maxContextconfiguration is too low, causing early conversation history to be truncated and failing to effectively convey contextual information. - Responses involving numerical information like dosage or time lack units or have incorrect numerical precision. This often happens because prompts do not explicitly require the model to focus on and output complete numerical values and units, or the original knowledge base text represents numerical values inconsistently.
How to Confirm Proper Configuration
- Conduct multi-turn conversation tests on core regulatory clauses to verify if the model accurately understands and coherently answers upstream and downstream questions.
- Randomly select specialized terms from documents and ask and follow-up questions in multi-turn conversations to confirm the model's understanding and correct use of these terms.
- Select procedural questions containing specific numerical values and units. Test the completeness and accuracy of numerical values and units in the model's responses. For example, when asked about "the dosage of a certain drug," the model should provide a numerical value with units.
The values provided are common starting points and should be measured against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.