Context and Tokens for Pharmacovigilance R&D Document Analysis

Pharmacovigilance data primarily comes from post-marketing adverse event reports, clinical study reports, drug label updates, and regulatory safety

Data Characteristics in Pharmacovigilance

Pharmacovigilance data primarily comes from post-marketing adverse event reports, clinical study reports, drug label updates, and regulatory safety information. These documents are typically semi-structured or unstructured text. Examples include PDF case report forms, Word documents of expert evaluations, or database records with free-text descriptions. Update frequency varies based on drug lifecycle and regulatory requirements. Updates can be quarterly, annually, or real-time for serious adverse events.

Document structures are diverse. They may include patient demographics, medication history, adverse event descriptions, diagnostic results, and treatment measures. Key specific information includes medical terminology, drug names, dosage units (e.g., mg/kg), and time units (e.g., days, weeks).

Constraints on Context and Tokens

The semi-structured nature of pharmacovigilance documents makes fixed-length chunking ineffective for capturing complete semantic information. For example, descriptions of causality between adverse events and related drugs may span multiple paragraphs. High frequencies of medical terms and specialized abbreviations increase tokenization complexity. This can lead to inaccurate tokenization or splitting of important entities.

Documents often contain large amounts of non-critical information, such as administrative headers, footers, or lengthy background introductions. This information unnecessarily consumes context token limits. Focus on time-series events requires the context to effectively link symptoms and medication use at different time points. This poses challenges for sliding window or multi-turn dialogue context management.

Processing these documents requires balancing information granularity with context length. Ensure the completeness of critical facts while preventing irrelevant information from interfering with model inference.

Configuration Settings

Configuration ItemRecommended ValueRationale
Chunk size (Chunk Length)1000–1500 charactersBalances the completeness of adverse event descriptions with model processing efficiency.
Chunk overlap (Chunk Overlap)100–200 charactersEnsures key information across chunks, such as events and drugs, can be effectively linked.
maxContext3500–4000 tokensMost models perform stably within this range, covering critical event chains.
Recall count (Recall Count)5–8 entriesBalances relevance with token consumption, reducing interference from non-critical information.
Similarity threshold (Similarity Threshold)0.75–0.82Focuses on highly relevant medical events and drug information.
Rerank result count (Reranked Return Count)3–5 entriesFurther refines results, increasing the weight of core information.

Common Mistakes

  • Retrieval results contain many non-core medical terms or irrelevant background information. This happens when the Similarity threshold (Similarity Threshold) is set too low, recalling overly generalized data blocks.
  • The model cannot accurately answer the complete causal chain of an adverse event. This manifests as a lack of critical time or dosage information in the answer. The Chunk size (Chunk Length) may be too short, or Chunk overlap (Chunk Overlap) insufficient, causing important information to be fragmented during chunking.
  • The model loses context for subsequent questions after the workflow clears the context, even when the user did not explicitly clear it. This can occur if maxContext or max_tokens parameters are set incorrectly, leading to automatic context truncation when the limit is reached.

How to Verify Configuration

  • Run a series of queries containing typical adverse event reports. Check if the Recall count (Recall Count) and Rerank result count (Reranked Return Count) focus on core drug, symptom, and time information. Adjust the Similarity threshold (Similarity Threshold) accordingly.
  • For complex adverse events described across multiple paragraphs, test if the model can fully extract and integrate all relevant information. This verifies the reasonableness of Chunk size (Chunk Length) and Chunk overlap (Chunk Overlap).
  • Monitor FastGPT's backend token usage statistics (e.g., through countGptMes related logs or metrics). Ensure that token consumption for each Q&A session remains within the maxContext limit and that no truncation warnings appear due to excessive context length.

Note: The values provided are common starting points. Measure them against your own samples for optimal performance.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.