Multiturn Conversation and Prompts for Indication-Based Drug Use Q&A

Indication data primarily comes from drug inserts, clinical guidelines, national drug administration databases, and professional medical literature.

Data Characteristics for This Category

Indication data primarily comes from drug inserts, clinical guidelines, national drug administration databases, and professional medical literature. This data typically exists as unstructured text, semi-structured tables, or structured fields. Update frequency varies: drug inserts and clinical guidelines are revised periodically due to new drug approvals, clinical research advancements, or regulatory policy changes, usually quarterly or annually. Document structures include standard sections like [Indications], [Dosage and Administration], and [Adverse Reactions] in drug inserts. Clinical guidelines focus more on disease diagnosis, treatment plans, and recommendation levels. Indication fields often involve disease names, disease codes (e.g., ICD-10), drug names, approval status, and restrictions for specific populations (e.g., children, pregnant women). Units are typically text descriptions and do not involve complex numerical units.

Constraints Imposed by These Characteristics on Multiturn Conversation and Prompts

The periodic updates of indication data require FastGPT's knowledge base to undergo regular incremental updates or re-indexing to ensure the accuracy of conversation content. The coexistence of unstructured and semi-structured data means that knowledge base construction must handle both text segmentation and table extraction to avoid information loss. For example, a single drug may be mentioned across multiple indications, with each indication potentially having different dosages or precautions. In multiturn conversations, users might inquire about an indication, then a related drug, and then its adverse reactions. This demands that the dialogue system maintain contextual coherence and accurately extract information from different knowledge fragments. Prompt design must guide the model to identify disease entities and drug entities in user queries, and to distinguish whether the query intent is to obtain a list of indications, specific indication details, or drug use guidance related to an indication.

Configuration Settings

Configuration ItemRecommended ValueRationale for Recommendation
maxContext6Indication-related queries often involve multiple follow-up questions, requiring a longer conversation history to ensure contextual coherence.
Chunk size (Segment Length)400–600 charactersIndication descriptions in drug inserts and clinical guidelines are often lengthy; this prevents key information from being truncated.
Recall count (Recall Count)10 entriesImproves the comprehensiveness of initial recall, covering multiple potentially relevant indications or drug information.
Similarity threshold (Similarity Threshold)0.75Ensures the precision of recalled content, filtering out document segments with low relevance to the user's query.
Rerank result count (Reranked Return Count)3 entriesIn multiturn conversations, users typically focus on the most direct answers, reducing redundant information.
RAG_MAX_TOKEN2000Background information and descriptions for indications can be long; this ensures the full context is sent to the large model for inference.

Common Mistakes

  • Phenomenon: FastGPT cannot accurately distinguish whether the user is asking about the indications for Drug A or Drug B in a multiturn conversation. Reason: The prompt fails to effectively guide the model to identify drug entities in the user's intent, or the knowledge base's segmentation strategy mixes indication information for different drugs.
  • Phenomenon: When a user asks for detailed information about an indication, the returned answer contains a large amount of irrelevant content, or states "cannot read this file." Reason: The knowledge base failed to correctly parse the table structure when processing the uploaded XLSX file, preventing indication information from being effectively indexed.
  • Phenomenon: When a user follows up with a question about the "dosage and administration" or "adverse reactions" for an indication, the model cannot provide the corresponding information or gives an answer inconsistent with the previous turn. Reason: The maxContext parameter is set too low, causing the conversation history context to be truncated and the model to lose memory of the indication from the previous turn.

How to Confirm Correct Configuration

  • Conduct multiturn conversation tests, starting with a specific disease name, then progressively asking about corresponding drugs, dosages, and adverse reactions. Verify that the conversation flow is smooth and the information is accurate and coherent.
  • Upload new drug inserts or clinical guideline files. Observe whether the knowledge base correctly parses and indexes the indication information within them. Then, try asking questions through dialogue to verify the availability of this new content.
  • Design ambiguous or context-dependent queries, such as "What are the indications for that drug?" Verify whether the model can accurately identify "that drug" based on the conversation history and provide the corresponding indications.

The values provided are common starting points and should be measured against specific samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.