Multi-turn Conversation and Prompts for Solid Tumor Pharmacovigilance

Solid tumor pharmacovigilance data comes from clinical trial reports, real-world studies, adverse event reporting systems (e.g., CIOMS reports)

Data Characteristics

Solid tumor pharmacovigilance data comes from clinical trial reports, real-world studies, adverse event reporting systems (e.g., CIOMS reports), medical literature, and drug inserts. Data update frequencies vary. Clinical trial data typically releases centrally after trials conclude, while real-world data may collect continuously and update periodically. Document structures are diverse, including structured Case Report Forms (CRFs), semi-structured medical narratives, and unstructured free-text reports. Core fields include patient demographics, diagnostic information, medication history, adverse event descriptions (MedDRA codes), event onset time, severity, outcome, and drug-relatedness assessments. Units involve dosage (mg, g), frequency (times/day), and duration (days, weeks, months).

Constraints on Multi-turn Conversations and Prompts

The multi-source and heterogeneous nature of solid tumor pharmacovigilance data challenges information integration in multi-turn conversations. Unstructured text reports contain numerous medical terms, abbreviations, and colloquialisms, requiring robust semantic understanding. Non-real-time data updates require models to process historical data and identify information timeliness. Adverse event severity and relatedness assessments often involve complex, multi-dimensional judgments. Prompt design must guide the model in logical reasoning. Inconsistent field names and units across different data sources require standardization during preprocessing to avoid ambiguity in conversations. Multi-turn conversations need context tracking to ensure coherent questioning and summarization of adverse events for the same patient or drug.

Configuration Settings

Configuration ItemSuggested ValueRationale
maxContext8Ensures coverage of 4-6 key conversation turns when querying adverse event details.
Chunk size (Segment Length)500–800 characters (characters)Accommodates long sentences and detailed descriptions in medical reports while preventing individual segments from being too long and affecting recall efficiency.
Recall count (Recall Count)15–20 entries (items)Increases recall coverage, considering adverse event reports may involve multiple related factors.
Similarity threshold (Similarity Threshold)0.75–0.85Balances recall precision and recall rate, avoiding interference from irrelevant information while not missing potentially relevant reports.
Rerank result count (Reranked Return Count)5 entries (items)Selects the most relevant core information from recalled results to reduce model processing burden and focus on key points.
PARSE_FILE_TIMEOUT_SECONDS300 seconds (seconds)Most solid tumor clinical report files are large, requiring sufficient time for parsing to prevent upload failures.

Common Pitfalls

  • Missing key field information in conversation output: The prompt did not explicitly instruct the model to extract and display specific fields; the model provided only general answers.
  • Inconsistent assessment of adverse event severity or relatedness after multi-turn conversations: Poor context management caused the model to forget previous judgments in later turns.
  • System unresponsiveness or errors after uploading large medical report files: PARSE_FILE_TIMEOUT_SECONDS or UPLOAD_FILE_MAX_SIZE parameters are typically set too low, causing file parsing or upload timeouts.

How to Verify Configuration

  • Conduct at least 5 simulated conversations for typical adverse events across different solid tumor types. Check if the model accurately identifies adverse event names, severity, and related drugs.
  • Randomly select 10 clinical adverse event reports. Use knowledge base Q&A to verify if the model correctly extracts core information from the reports. Compare with original reports; the error rate should be below a preset threshold.
  • Upload a free-text report containing complex medical terms and abbreviations. Test if the model can correctly interpret these terms in multi-turn conversations and maintain semantic consistency.

Note: The values provided are common starting points. Measure them against your own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.