Multi-Turn Conversations and Prompts for mRNA Vaccine Pharmacovigilance

mRNA vaccine pharmacovigilance data comes from clinical trial reports, real-world evidence (RWE), patient voluntary reporting systems (e.g., VAERS

Data Characteristics

mRNA vaccine pharmacovigilance data comes from clinical trial reports, real-world evidence (RWE), patient voluntary reporting systems (e.g., VAERS, EudraVigilance), and medical literature. This data mixes unstructured text (e.g., patient descriptions, physician diagnoses) and structured data (e.g., adverse event codes, occurrence dates, dosages, batch numbers). Data updates frequently, especially during initial vaccine release and mass vaccination campaigns, with new reports generated daily or even hourly. Document structures vary, containing medical terminology, abbreviations, dosage units (e.g., μg), time units (e.g., days, hours), and descriptions of individual patient differences. Adverse event (AE) severity and causality assessments are often qualitative, requiring expert interpretation.

Constraints on Multi-Turn Conversations and Prompts

The real-time nature and high update frequency of mRNA vaccine adverse event reports require multi-turn conversation systems to quickly integrate the latest knowledge, avoiding information lag. The prevalence of unstructured text makes natural language understanding critical, especially for accurate recognition of medical terminology and colloquial patient descriptions. Multi-turn conversations need to support follow-up questions and clarifications on key fields in adverse event reports (e.g., 疫苗批次, 接种日期, 症状, Duration time) to ensure data completeness. Standardized handling of dosage and time units is crucial for avoiding ambiguity and subsequent analysis. The complexity of causality assessment requires prompt design to guide the model in providing relevant information and possibility analyses without giving definitive diagnoses, avoiding misinformation.

Configuration Settings

Configuration ItemRecommended ValueRationale
maxContext8 turnsMaintains conversational coherence and controls token consumption. maxContext set to 8 turns covers most adverse event reporting scenarios.
Chunk size500 charactersBalances text semantic integrity and retrieval efficiency. A segment length of 500 characters is suitable for medical texts.
Recall count7 itemsIncreases recall coverage to obtain more comprehensive relevant knowledge, addressing multi-source data.
Similarity threshold0.78Balances recall accuracy and breadth, avoiding interference from irrelevant information. 0.78 has been empirically shown to work well for medical texts.
Rerank result count3 itemsHighlights the most relevant knowledge snippets, reducing model processing burden and focusing on core information.
promptTemplateIncludes adverse events Reporting GuidelinesGuides the model to adhere to pharmacovigilance guidelines when answering, for example, emphasizing information sources and caution in causality assessment.

Common Pitfalls

  • Conversation results frequently show "insufficient information to judge": This occurs when key entities in adverse event reports (e.g., drug name, 不良反应症状, Occurrence Time) are extracted inaccurately, leading to knowledge base retrieval failures or irrelevant retrieved content.
  • The model "forgets" previously mentioned information in multi-turn conversations: This happens when the maxContext parameter is set too low, causing historical conversation information to be truncated from the context window.
  • HTTP 401 Unauthorized errors in prompts: This indicates incorrect or expired API key configuration, preventing FastGPT from calling external model interfaces.

Verification Steps

  • Simulate multi-turn conversations to check if the system accurately identifies and asks follow-up questions about core information in adverse event reports, such as 接种日期, 症状描述, 疫苗批次号.
  • Check log output to confirm that the Recall count and Similarity threshold for knowledge base retrieval meet expectations, and that retrieved knowledge snippets are highly relevant to the current conversation.
  • Verify that the system correctly understands and provides relevant information when processing prompts containing medical abbreviations and dosage units (e.g., mg, μg, ml).

The values provided are common starting points and should be measured against the reader's own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.