Data Characteristics in Pharmacovigilance
Pharmacovigilance data primarily comes from post-marketing adverse event reports, clinical study reports, drug label updates, and regulatory safety information. These documents are typically semi-structured or unstructured text. Examples include PDF case report forms, Word documents of expert evaluations, or database records with free-text descriptions. Update frequency varies based on drug lifecycle and regulatory requirements. Updates can be quarterly, annually, or real-time for serious adverse events.
Document structures are diverse. They may include patient demographics, medication history, adverse event descriptions, diagnostic results, and treatment measures. Key specific information includes medical terminology, drug names, dosage units (e.g., mg/kg), and time units (e.g., days, weeks).
Constraints on Context and Tokens
The semi-structured nature of pharmacovigilance documents makes fixed-length chunking ineffective for capturing complete semantic information. For example, descriptions of causality between adverse events and related drugs may span multiple paragraphs. High frequencies of medical terms and specialized abbreviations increase tokenization complexity. This can lead to inaccurate tokenization or splitting of important entities.
Documents often contain large amounts of non-critical information, such as administrative headers, footers, or lengthy background introductions. This information unnecessarily consumes context token limits. Focus on time-series events requires the context to effectively link symptoms and medication use at different time points. This poses challenges for sliding window or multi-turn dialogue context management.
Processing these documents requires balancing information granularity with context length. Ensure the completeness of critical facts while preventing irrelevant information from interfering with model inference.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Chunk Length) | 1000–1500 characters | Balances the completeness of adverse event descriptions with model processing efficiency. |
Chunk overlap (Chunk Overlap) | 100–200 characters | Ensures key information across chunks, such as events and drugs, can be effectively linked. |
maxContext | 3500–4000 tokens | Most models perform stably within this range, covering critical event chains. |
Recall count (Recall Count) | 5–8 entries | Balances relevance with token consumption, reducing interference from non-critical information. |
Similarity threshold (Similarity Threshold) | 0.75–0.82 | Focuses on highly relevant medical events and drug information. |
Rerank result count (Reranked Return Count) | 3–5 entries | Further refines results, increasing the weight of core information. |
Common Mistakes
- Retrieval results contain many non-core medical terms or irrelevant background information. This happens when the
Similarity threshold(Similarity Threshold) is set too low, recalling overly generalized data blocks. - The model cannot accurately answer the complete causal chain of an adverse event. This manifests as a lack of critical time or dosage information in the answer. The
Chunk size(Chunk Length) may be too short, orChunk overlap(Chunk Overlap) insufficient, causing important information to be fragmented during chunking. - The model loses context for subsequent questions after the workflow clears the context, even when the user did not explicitly clear it. This can occur if
maxContextormax_tokensparameters are set incorrectly, leading to automatic context truncation when the limit is reached.
How to Verify Configuration
- Run a series of queries containing typical adverse event reports. Check if the
Recall count(Recall Count) andRerank result count(Reranked Return Count) focus on core drug, symptom, and time information. Adjust theSimilarity threshold(Similarity Threshold) accordingly. - For complex adverse events described across multiple paragraphs, test if the model can fully extract and integrate all relevant information. This verifies the reasonableness of
Chunk size(Chunk Length) andChunk overlap(Chunk Overlap). - Monitor FastGPT's backend token usage statistics (e.g., through
countGptMesrelated logs or metrics). Ensure that token consumption for each Q&A session remains within themaxContextlimit and that no truncation warnings appear due to excessive context length.
Note: The values provided are common starting points. Measure them against your own samples for optimal performance.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.