Vector Models and Indexing for Hematologic Oncology Pharmacovigilance

Hematologic oncology pharmacovigilance data primarily originates from clinical trial reports, real-world evidence (RWE) studies, case reports, medical

Data Characteristics

Hematologic oncology pharmacovigilance data primarily originates from clinical trial reports, real-world evidence (RWE) studies, case reports, medical literature, and drug regulatory safety updates. This data is largely unstructured text, such as patient medical records and free-text descriptions in Individual Case Safety Reports (ICSRs). Structured data includes drug names, adverse reaction terms (MedDRA codes), dosages, treatment regimens, and patient demographic information. Data updates frequently, especially after new drug launches or when new safety signals emerge. Documents often contain extensive specialized medical terminology, abbreviations, and critical time-series information (e.g., adverse event onset and duration).

Constraints on Vector Models and Indexing

The heterogeneous nature of hematologic oncology pharmacovigilance data requires vector models to effectively integrate text and structured information. High update frequency necessitates an indexing mechanism that supports efficient incremental updates to ensure timely retrieval. Specialized medical terminology, abbreviations, and complex causal relationships among disease progression, treatment regimens, and adverse reactions demand advanced semantic understanding from vector models. Generic vector models may struggle to capture these subtle medical semantics, leading to reduced retrieval accuracy. Furthermore, adverse event reports are often lengthy and contain multiple entities and events, requiring fine-grained text segmentation strategies to prevent information loss or over-generalization. The strong correlation of time-series information also requires vector index design to consider the temporal dimension, aiding time-series-related retrieval.

Configuration Guidelines

Configuration ItemRecommended ValueRationale
Chunk size (Segment Length)300-500 charactersBalances completeness of adverse event descriptions with vectorization efficiency.
Chunk Overlap Length (Segment Overlap Length)50-100 charactersEnsures context continuity and prevents important information from being split.
embedding_modelVertical-domain fine-tuned modelImproves understanding of medical terminology and complex semantics.
Similarity threshold (Similarity Threshold)0.75-0.85Reduces false positives and recalls highly relevant adverse events.
Recall count (Recall Count)20-30 entriesProvides sufficient candidate results for subsequent re-ranking and analysis.
Max Index File Size500 MBBalances index build speed and retrieval performance, avoiding overly large single files.

Common Pitfalls

  • Connecting to VLLM-deployed Embedding models may result in connection timeouts or authentication failures. This typically stems from network configuration issues, incorrect API keys, or the model service not starting correctly.
  • Knowledge base retrieval results may deviate significantly from expectations. A common cause is inappropriate text segmentation granularity, leading to key information being split or insufficient context for meaningful semantic blocks.
  • Locally deployed FastGPT may experience abnormal disk usage growth. This can relate to ineffective cleanup of original documents, segmented text blocks, and vector data, or index fragmentation.

Verification Steps

  • Check the connection status of the embedding_model through the FastGPT administration interface. Ensure it returns a 200 status code.
  • Upload a typical hematologic oncology adverse event report. Observe the knowledge base segmentation preview to confirm Chunk size and Chunk Overlap Length meet expectations.
  • Perform a retrieval query using specific medical terminology. Evaluate if the returned results include highly relevant hematologic oncology adverse event reports and check the filtering effect of the Similarity threshold.
  • Regularly monitor FastGPT server storage usage. Compare actual consumption with the expected growth rate.

The values provided are common starting points. Measure them against your own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.