Multi-Turn Conversations and Prompts for Solid Tumor Regulatory Submission Document Preparation

Solid tumor regulatory submission documents primarily include clinical trial protocols, clinical study reports (CSRs), drug labels, pharmaceutical

Data Characteristics for this Category

Solid tumor regulatory submission documents primarily include clinical trial protocols, clinical study reports (CSRs), drug labels, pharmaceutical research reports, and non-clinical study reports. Data sources are extensive, involving hospital medical record systems, laboratory test data, biostatistical reports, and medical imaging data. Update frequency can be high, influenced by clinical trial progress, changes in approval policies, and new research findings. Document structures are typically chapter-based, following regulatory guidelines such as ICH GCP and NMPA. They contain numerous standardized medical terms, biomarker names, dosage units (e.g., mg/kg, mg/m²), and time units (e.g., weeks, months). Field specificity is high, for example, RECIST 1.1 evaluation criteria, PD-L1 expression levels, and clinical endpoint data like ORR and PFS.

Constraints from these Characteristics on Multi-Turn Conversations and Prompts

The specialized nature, standardization, and update frequency of solid tumor submission documents impose specific requirements on multi-turn conversation context management and prompt design. The vast number of specialized terms and abbreviations requires strong semantic understanding to prevent misunderstandings due to terminological ambiguity in multi-turn conversations. The chapter-based document structure and internal references necessitate deep understanding of document interrelations to support cross-chapter information retrieval. High-frequency data updates, especially for clinical guidelines and regulatory policies, require the knowledge base to synchronize quickly, ensuring conversations are based on the latest information. Precision for units like dosage and time requires prompts to guide the model in strict numerical extraction and comparison, avoiding errors due to unit confusion. The long document characteristic makes context window management crucial for maintaining conversation coherence.

Configuration Settings

Configuration ItemSuggested ValueRationale for this Value
maxContext6000–8000 charactersSolid tumor submission documents are long. A larger context window maintains multi-turn conversation coherence and reduces information loss.
Recall Count8–12 itemsEnsures enough relevant snippets are recalled from vast professional documents, covering complex medical concepts and related information.
Similarity Threshold0.78–0.85Balances precision and recall. Avoids missing critical specialized term matches with too high a threshold, or introducing irrelevant information with too low a threshold.
Segment Length500–800 charactersConsiders the semantic integrity of medical text. Avoids truncating key information snippets, which affects understanding.
Rerank Return Count5 itemsFurther refines the most relevant snippets after initial recall, improving accuracy in multi-turn conversations.
QUERY_REWRITE_MODESmart RewriteFor complex multi-turn questions and specialized terms, smart rewriting better understands user intent and generates more precise queries.

Three Common Mistakes

  • Slow conversation response speed, sometimes with an Unexpected end of JSON input error. This may be due to insufficient computational resources when the model processes long contexts or complex queries, or the backend service terminating before returning a complete JSON response.
  • The model "forgets" in multi-turn conversations, failing to link to solid tumor clinical data from previous turns. This usually happens when the maxContext parameter is set too small, causing historical conversation information to be truncated and not retained in the current context window.
  • Confusion of dosage or time units in responses, such as misinterpreting mg/kg as mg/m², or weeks as months. This occurs when prompts do not explicitly guide the model to focus on and verify numerical units, or when unit standardization in the knowledge base for relevant fields is insufficient.

How to Confirm Proper Configuration

  • Select a solid tumor submission document with complex clinical trial data and pharmaceutical information. Conduct multi-turn conversation tests to check if the model accurately understands and links specialized terms and data across different chapters, and evaluate conversation coherence.
  • Randomly select 10 queries containing dosage or time units. Verify if the model's extraction and use of these units in responses are accurate, and compare with original documents to confirm no unit confusion.
  • Simulate user questions about the latest clinical guidelines or approval policies. Check if the model can cite the most recently updated document content from the knowledge base and provide the correct regulatory requirement version number.

Note: The values provided are common starting points. Measure them against your own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.