Multi-turn Conversation and Prompts for Infectious Disease Pharmacovigilance

Infectious disease pharmacovigilance data primarily originates from national drug adverse reaction monitoring centers, medical institution case

Data Characteristics for this Category

Infectious disease pharmacovigilance data primarily originates from national drug adverse reaction monitoring centers, medical institution case reports, clinical trial reports, and relevant academic literature. Data updates are frequent, typically summarized and released quarterly or monthly, with some serious adverse event reports updated in real-time. Document structures are predominantly a mix of structured reports and unstructured text. Structured sections include fields such as patient basic information, drug information (e.g., ATC_CODE, BATCH_NUMBER), adverse event details (AE_TERM, SERIOUSNESS_CRITERIA), and outcomes. Unstructured sections provide detailed case descriptions, potentially covering disease progression, concomitant medications, laboratory test results, and healthcare professional observations. Units commonly used for dosage are mg, g, IU; frequency often uses times/day; and duration typically uses days, weeks, months.

Constraints Imposed by these Characteristics on Multi-turn Conversation and Prompts

The high update frequency of infectious disease pharmacovigilance data requires the knowledge base to have rapid synchronization and update mechanisms. This ensures the timeliness of information referenced in multi-turn conversations. The coexistence of structured and unstructured data necessitates prompt design that balances precise queries for specific fields with semantic understanding of free text. For example, a user might directly ask about "phototoxicity caused by a certain antibiotic," which requires extracting information from unstructured descriptions. The specialized terminology for adverse events (e.g., AE_TERM) and standardized dosage units demand that prompts parse user input by mapping colloquial expressions to professional terms and standard units. This avoids retrieval failures due to inconsistent terminology. Furthermore, understanding the context of patient medical history and medication history in multi-turn conversations is crucial for assessing the causality of adverse reactions.

Configuration Settings

Configuration ItemRecommended ValueRationale
maxContext2048 tokenEnsures sufficient context is retained when processing complex case descriptions, accommodating multi-turn follow-up questions.
Chunk size (Segment Length)500 characters (characters)Balances semantic completeness with recall efficiency, preventing long paragraphs from diluting key information.
Recall count (Recall Count)Top 8 entries (top 8)The number of infectious disease adverse event reports is large; increasing the recall count can improve relevance.
Similarity threshold (Similarity Threshold)0.75Balances recall rate and accuracy, filtering out low-relevance document segments.
Rerank result count (Rerank Return Count)Top 3 entries (top 3)After reranking, focuses on the most relevant few pieces of information, reducing the model's inference burden.
PARSE_FILE_TIMEOUT_SECONDS600 seconds (seconds)Provides ample parsing time when processing large PDF files or structurally complex case reports.

Three Common Mistakes

  • After uploading a case file in a conversation, the system returns a "file parsing timeout" error. This occurs because the PARSE_FILE_TIMEOUT_SECONDS parameter is set too low, failing to cover the parsing duration for some complex or large files.
  • A user asks about the "incidence of liver injury for a certain antiviral drug," but the conversational AI provides a general answer like "may cause liver injury" instead of accurate data. This happens because the prompt does not explicitly guide the model to extract specific numerical values from structured data fields (e.g., ADVERSE_EVENT_RATE), relying instead on general text matching.
  • In a multi-turn conversation, the user repeatedly mentions the same pathogen infection, but the system fails to effectively utilize this context in subsequent replies, leading to repeated questions or information omission. This indicates that the conversation management logic does not fully leverage the maxContext parameter or that prompts are not effectively designed to maintain conversational focus.

How to Confirm Proper Configuration

  • Select typical case reports and simulate multi-turn conversations. Verify that the system accurately extracts patient chief complaints, medication history, and adverse events, and continuously tracks key information.
  • Upload infectious disease-related literature or case reports of different formats and sizes. Observe if file parsing is successful and check if parsing time is within the expected range. This confirms the PARSE_FILE_TIMEOUT_SECONDS setting.
  • Conduct query tests for specific drug and adverse event combinations. Compare the system's returned results with the original knowledge base content. This ensures that Similarity threshold (Similarity Threshold) and Recall count (Recall Count) effectively retrieve relevant and accurate information.

The values provided are common starting points. Measure them against your own samples to determine optimal settings.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.