Data Characteristics for This Category
Phase I clinical pharmacovigilance data originates from clinical observation records, adverse event reports, laboratory test results, and concomitant medication information collected from subjects during trials. This data typically exists as a mix of structured and unstructured formats. Structured data includes adverse event (AE) codes (e.g., MedDRA terms), adverse reaction onset times, severity, and outcomes from electronic health record (EHR/EDC) systems. Unstructured data consists of free-text descriptions from doctors or nurses, detailing reported symptoms, signs, physician assessments, and causality judgments related to the investigational drug. Data updates frequently, often within hours or days of an adverse event. Document types are diverse, encompassing case report forms (CRFs), medical imaging reports, pathology reports, and informed consent forms. Field units vary; for example, laboratory indicators include mg/dL and mmol/L, and adverse event onset times are precise to hours or even minutes.
Constraints Imposed by These Characteristics on "Multi-turn Conversations and Prompts"
The multi-source and heterogeneous nature of Phase I clinical pharmacovigilance data demands high robustness from multi-turn dialogue systems when understanding and integrating information. High update frequency requires near real-time synchronization of the knowledge base to ensure the dialogue model accesses the latest information. The richness of unstructured free text challenges prompt engineering, necessitating finely designed prompts to guide the model in accurately extracting key information. For example, distinguishing between "adverse event occurrence" and "adverse event reporting time." The specialized nature of fields and units requires the model to accurately cite and explain these professional terms in its responses, avoiding misinterpretation. For instance, the model must understand the normal ranges and clinical significance of different laboratory indicators. Furthermore, subject privacy and data sensitivity mandate strict adherence to data security and compliance standards when processing and displaying information, preventing sensitive information leakage. This impacts retrieval strategies and content filtering for generation.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale for Recommendation |
|---|---|---|
maxContext | 6 turns | Balances short-term memory with information volume, preventing model confusion or excessive computational resource consumption from overly long contexts. |
Chunk size (Segment Length) | 500 characters | Accommodates the length of descriptive text in clinical reports, ensuring the completeness of single-segment information. |
Recall count (Retrieval Count) | top 8 items | Ensures coverage of multiple relevant adverse event reports or subject records, enhancing information comprehensiveness. |
Similarity threshold (Similarity Threshold) | 0.75 | Balances retrieval precision and coverage, filtering irrelevant information, and reducing noise. |
Rerank result count (Reranked Return Count) | top 3 items | Focuses on the most relevant key information, reduces the model's processing burden, and improves response accuracy. |
TEMPERATURE | 0.3 | Controls the certainty of model output, ensuring the rigor and accuracy of pharmacovigilance information, and preventing hallucinations. |
Three Common Mistakes
- The dialogue model includes adverse event information unrelated to the current subject in its answer. This occurs when the
Similarity threshold(Similarity Threshold) in knowledge base retrieval is set too low, leading to the retrieval of irrelevant documents and the model generating content based on incorrect context. - The model fails to understand certain laboratory indicator abbreviations or units, leading to deviations in response content. This happens when the knowledge base lacks standardized mapping or explanations for professional terms and units, and prompts fail to effectively guide the model in professional knowledge retrieval.
- In multi-turn conversations, the model's understanding of user queries drifts, failing to continuously focus on the same subject or adverse event. This is due to
maxContextbeing set too short, causing the model to lose early dialogue context information.
How to Confirm Proper Configuration
- Simulate multi-turn conversations for typical adverse event scenarios. Verify the model's accuracy in extracting and understanding key information such as subject basic information, adverse event descriptions, and laboratory indicators.
- Check that professional terms and data units cited by the model in its responses are consistent with the original knowledge base text, without confusion or alteration.
- By tracking FastGPT's dialogue logs, analyze the content of documents retrieved by the knowledge base across different dialogue turns. Confirm that the retrieval count and similarity scores meet expectations, and no irrelevant information is retrieved.
Note: The values provided are common starting points. Measure them against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.