Multi-Turn Conversations and Prompts for Rare Disease Products

Rare disease product data comes from diverse sources. These include global rare disease databases (e.g., ORPHANET, OMIM), clinical trial reports, drug

Data Characteristics in This Category

Rare disease product data comes from diverse sources. These include global rare disease databases (e.g., ORPHANET, OMIM), clinical trial reports, drug labels, medical literature, and patient registry data. Update frequencies vary. Drug labels and clinical trial data typically update with approvals and research progress. Patient registry data accumulates continuously. Document structures often include highly structured tabular data (e.g., gene mutation sites, indications, dosage, adverse reactions) and semi-structured or unstructured text descriptions (e.g., disease mechanisms, diagnostic criteria, treatment plans). Fields and units are highly specialized. For example, gene sequences use base pair units, drug dosages often express in mg/kg body weight or units/day, and disease progression indicators may involve specific biomarker concentrations or scale scores.

Constraints on Multi-Turn Conversations and Prompts Due to These Characteristics

The highly specialized and diverse nature of rare disease data imposes specific requirements on the accuracy of multi-turn conversations and prompt construction. Due to varying update frequencies, the system must identify and prioritize the latest, most authoritative data sources. This prevents the use of outdated information. Complex document structures mean that RAG (Retrieval-Augmented Generation) mechanisms require refined processing. The system must accurately extract numerical values from structured data and summarize information from unstructured text. The specialized nature of fields and units requires prompts to strictly adhere to medical terminology and measurement unit specifications when generating responses. This avoids ambiguity or errors. In multi-turn conversations, users may progressively ask more in-depth questions, for example, from disease overview to specific gene therapy protocols. This requires the system to maintain contextual coherence and dynamically adjust retrieval strategies and information organization based on previous conversations.

Configuration Settings

Configuration ItemRecommended ValueRationale for This Value
maxContext2048 tokensBalances contextual coherence with model processing efficiency, accommodating complex rare disease descriptions.
Recall countTop 5 entriesEnsures relevance of retrieval results, reducing interference from unnecessary information.
Similarity threshold0.78Ensures retrieved text segments are highly relevant to the user query, filtering noise.
Chunk size500 charactersOptimizes text segmentation granularity, preventing a single segment from containing too much irrelevant information.
Rerank result countTop 3 entriesFurther refines retrieval results, providing the most core reference information.
promptCalibrate by actual measurementIteratively optimize based on rare disease terminology and query patterns.

Three Common Mistakes

  1. Conversation error showing 400 status code (no body): This usually occurs when the prompt is too long or contains illegal characters, causing the request body to not conform to API specifications.
  2. Responses contain dosage unit confusion or incorrect gene locus descriptions: This happens because RAG-retrieved text segments fail to effectively distinguish synonyms or similar units, and the prompt does not explicitly emphasize unit and format rigor.
  3. Conversations fail to maintain context, frequently repeating questions or providing irrelevant information: This occurs when the maxContext parameter is set too low, or the context management mechanism fails to effectively handle entity references and information accumulation in multi-turn conversations.

How to Confirm Correct Configuration

  • Test if the system can accurately identify and cite the latest drug labels or clinical guidelines for a series of rare disease product queries.
  • Verify if the system can correctly track a user's progressively in-depth questions about specific genes, proteins, or disease mechanisms in multi-turn conversations and provide coherent responses.
  • Check if units and formats consistently comply with medical standards when responses involve professional information such as dosage, frequency, and gene mutation sites.
  • Simulate user queries to confirm if the system can proactively clarify or guide users to provide more key information when encountering vague or incomplete queries.

Note: The values provided are common starting points and should be measured against the reader's own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.