Data Characteristics for This Category
Pharmacovigilance data for academic promotion primarily originates from clinical trial reports, real-world evidence (RWE), case reports, drug labels, and various medical literature. Data update frequencies vary. Clinical trial data typically releases after study completion, while case reports and medical literature generate continuously. Document structures are diverse. These include structured database records (e.g., fields in adverse event reporting systems), semi-structured clinical study reports (containing abstracts, methods, results, discussions), and unstructured free-text descriptions (e.g., physician notes, patient interview records). Key fields include drug name, adverse event name, occurrence time, severity, outcome, patient demographics, and concomitant medications. For units, dosage commonly uses milligrams (mg) and grams (g). Time commonly uses days (day), weeks (week), and months (month). Severity often uses grading scales (e.g., CTCAE grades).
Constraints Imposed by These Characteristics on Multi-turn Conversations and Prompts
Diverse data sources require multi-turn dialogue systems to handle various data formats. Structured data allows direct matching. Unstructured text needs more complex semantic understanding and information extraction. Uncertain update frequencies mean knowledge base maintenance requires strategic planning to ensure dialogue content timeliness, especially for newly released adverse event information. Complex document structures pose challenges for prompt design. Prompts must guide the model to precisely locate key information from lengthy reports and avoid information redundancy. For example, when a user asks for detailed information on a specific adverse reaction to a drug, the system needs to integrate and extract relevant content from multiple sources. Standardization of fields and units is fundamental for accurate dialogue. Prompts need to clearly instruct the model to consistently use standardized terminology and units in responses, avoiding ambiguity.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
maxContext | 2000 characters | Balances dialogue coherence and computational resource consumption, ensuring coverage of key information across multiple turns. |
Chunk size | 500 characters | Accommodates paragraph lengths in medical literature, facilitating the model's understanding of complete semantics. |
Recall count | 8 entries | Balances information coverage and relevance, reducing interference from irrelevant information. |
Similarity threshold | 0.75 | Ensures recalled results are highly relevant to the user's query, filtering out vague matches. |
Rerank result count | 3 entries | Focuses on the most core information, improving response accuracy. |
temperature | 0.3 | Reduces model divergence, ensuring responses are fact-based and avoid hallucinations. |
Three Common Pitfalls
- Dialogue history loss after a dialogue window refresh, while backend records persist: This typically results from frontend caching issues or incorrect
sessionIdpassing. - AI dialogue nodes fail to retrieve file link variables: This often occurs because the
file_idorurlof an uploaded file is not correctly bound to context variables. - Model responses contain non-standardized medical terminology or units: This happens when prompts do not explicitly require the use of standard terminology sets, or the knowledge base contains non-standard data.
How to Confirm Correct Configuration
- Conduct multi-turn dialogue tests to verify the model's ability to maintain context coherence in complex scenarios and accurately cite knowledge base content.
- For different types of questions (e.g., adverse event details, drug comparisons, treatment recommendations), check if model responses are accurate, complete, and verify cited data sources.
- Simulate updates to newly released adverse event information. Test if the system can incorporate new information into dialogue responses within a reasonable timeframe.
The values provided are common starting points and should be measured against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.