Data Characteristics
Hematologic oncology pharmacovigilance data primarily originates from clinical trial reports, real-world evidence (RWE) studies, case reports, medical literature, and drug regulatory safety updates. This data is largely unstructured text, such as patient medical records and free-text descriptions in Individual Case Safety Reports (ICSRs). Structured data includes drug names, adverse reaction terms (MedDRA codes), dosages, treatment regimens, and patient demographic information. Data updates frequently, especially after new drug launches or when new safety signals emerge. Documents often contain extensive specialized medical terminology, abbreviations, and critical time-series information (e.g., adverse event onset and duration).
Constraints on Vector Models and Indexing
The heterogeneous nature of hematologic oncology pharmacovigilance data requires vector models to effectively integrate text and structured information. High update frequency necessitates an indexing mechanism that supports efficient incremental updates to ensure timely retrieval. Specialized medical terminology, abbreviations, and complex causal relationships among disease progression, treatment regimens, and adverse reactions demand advanced semantic understanding from vector models. Generic vector models may struggle to capture these subtle medical semantics, leading to reduced retrieval accuracy. Furthermore, adverse event reports are often lengthy and contain multiple entities and events, requiring fine-grained text segmentation strategies to prevent information loss or over-generalization. The strong correlation of time-series information also requires vector index design to consider the temporal dimension, aiding time-series-related retrieval.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Segment Length) | 300-500 characters | Balances completeness of adverse event descriptions with vectorization efficiency. |
Chunk Overlap Length (Segment Overlap Length) | 50-100 characters | Ensures context continuity and prevents important information from being split. |
embedding_model | Vertical-domain fine-tuned model | Improves understanding of medical terminology and complex semantics. |
Similarity threshold (Similarity Threshold) | 0.75-0.85 | Reduces false positives and recalls highly relevant adverse events. |
Recall count (Recall Count) | 20-30 entries | Provides sufficient candidate results for subsequent re-ranking and analysis. |
Max Index File Size | 500 MB | Balances index build speed and retrieval performance, avoiding overly large single files. |
Common Pitfalls
- Connecting to VLLM-deployed Embedding models may result in connection timeouts or authentication failures. This typically stems from network configuration issues, incorrect API keys, or the model service not starting correctly.
- Knowledge base retrieval results may deviate significantly from expectations. A common cause is inappropriate text segmentation granularity, leading to key information being split or insufficient context for meaningful semantic blocks.
- Locally deployed FastGPT may experience abnormal disk usage growth. This can relate to ineffective cleanup of original documents, segmented text blocks, and vector data, or index fragmentation.
Verification Steps
- Check the connection status of the
embedding_modelthrough the FastGPT administration interface. Ensure it returns a200status code. - Upload a typical hematologic oncology adverse event report. Observe the knowledge base segmentation preview to confirm
Chunk sizeandChunk Overlap Lengthmeet expectations. - Perform a retrieval query using specific medical terminology. Evaluate if the returned results include highly relevant hematologic oncology adverse event reports and check the filtering effect of the
Similarity threshold. - Regularly monitor FastGPT server storage usage. Compare actual consumption with the expected growth rate.
The values provided are common starting points. Measure them against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.