Data Characteristics
Site Management Organizations (SMOs) in pharmacovigilance primarily use data from clinical trial sites' electronic medical record systems, adverse event (AE/SAE) report forms, and subject interview records. This data exists in both structured formats (e.g., database records, XML files) and unstructured formats (e.g., PDF documents, scanned doctor's handwritten notes). Adverse event reports are immediate; serious adverse events (SAEs) typically require reporting within 24 hours, while general adverse events (AEs) are summarized periodically according to the study protocol. Adverse event reports follow international standards such as ICH E2B or FDA 3500A, including fields for patient demographics, drug information, adverse event description, diagnosis, management, and outcome.
Constraints Imposed on Multiturn Conversation and Prompts
SMO pharmacovigilance data characteristics impose specific constraints on building multiturn conversations and prompts. First, the immediacy and standardization requirements for adverse event reports mean the conversation system must respond quickly and accurately extract key information to support subsequent risk assessment and reporting. Multiturn conversations must identify and populate ICH E2B standard fields. Second, the mixed data sources (structured and unstructured) require stronger semantic understanding and information extraction accuracy when processing unstructured text. This avoids misinterpretations of critical medical terms and dosage units. For example, distinguishing between "mg/kg" and "mg" is crucial. Finally, the high update frequency requires the knowledge base to quickly synchronize the latest reports. Prompt design must guide the model to prioritize retrieving the newest data to ensure the timeliness of decision-making.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
maxContext | 8000 tokens | Accommodates a patient's complete medical history, medication details, and descriptions of multiple adverse events, providing comprehensive context. |
temperature | 0.3–0.5 | Balances accuracy and diversity in generated responses. Maintains rigor during key information extraction and provides appropriate flexibility during explanations. |
recall_top_k | 10 | Adverse event reports may involve multiple related factors and historical records. Increasing the number of recalled items comprehensively covers potentially associated information. |
Similarity threshold | 0.78–0.85 | Ensures recalled documents are highly relevant to the user's query. Avoids introducing irrelevant information that could interfere with adverse event judgment, especially for medical terminology. |
Chunk size | 500 characters | Balances textual semantic integrity and retrieval efficiency. Avoids overly long segments that dilute key information and overly short segments that lose context. |
prompt_template | Specific Instruction Set | Includes explicit instructions for ICH E2B field extraction, dosage unit identification rules (e.g., mg/kg, IU), and risk assessment logic guidance. |
Common Pitfalls
- Phenomenon: The model repeatedly asks for adverse event occurrence times or drug dosages already provided in the conversation. Reason: The prompt failed to effectively guide the model to identify and extract time or dosage unit fields, or context management did not correctly persist this information.
- Phenomenon: When processing a patient's medical history, the model confuses irrelevant past medical history with the current adverse event. Reason: The
Similarity threshold(similarity threshold) is set too low, leading to the recall of documents semantically distant from the current issue, or the prompt did not explicitly limit the focus to recent events. - Phenomenon: In a multiturn conversation, the model cannot accurately associate different symptoms of the same adverse event mentioned in different turns. Reason: The conversation management component failed to effectively track and merge mentions of the same entity across different turns, leading to fragmented information.
Verification of Configuration
- Simulate multiple conversation scenarios involving complex adverse event descriptions and historical medical records. Verify that the model accurately identifies and extracts key fields from ICH E2B standards in each turn, such as
patient_age,drug_name, andevent_onset_date. - Check if the model maintains unit and route accuracy when handling queries involving different dosage units (e.g.,
mg,g,IU) and administration routes (e.g.,oral,intravenous), and if it can correctly convert or explain them. - Compare summaries or risk assessments generated by the model for the same adverse event report in different conversation turns. Verify their consistency and logical coherence, ensuring no critical information is missed or contradictory.
- Assess if the model can effectively clarify medical abbreviations (e.g.,
BP,HR) or vague descriptions (e.g.,feeling unwell) through follow-up questions or by leveraging context, avoiding misjudgments.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.