Data Characteristics
Data in pharmacovigilance for DTP pharmacies originates from patient feedback, pharmacist follow-up records, drug sales records, and safety information from regulatory bodies. This data updates frequently. Patient feedback can be real-time, while pharmacist records are typically compiled daily or weekly. Document formats vary, including unstructured free text (e.g., patient descriptions of adverse reactions), structured tabular data (e.g., dosage, time, batch information), and semi-structured reports (e.g., adverse drug reaction report forms). Common fields include drug generic name, brand name, batch number, manufacturer, patient age, gender, underlying diseases, medication history, adverse reaction event description, occurrence time, duration, treatment measures, and outcome. Adverse reaction descriptions often contain medical terminology, colloquialisms, and even typos.
Constraints Imposed by Data Characteristics on Vector Models and Indexing
The diversity and high update frequency of DTP pharmacy data impose specific requirements on vector models and indexing. Unstructured text requires efficient text vectorization to capture semantic information. Structured fields need additional processing to integrate into the vector space or serve as metadata filters during retrieval. High update frequency necessitates incremental indexing and real-time updates for the knowledge base to ensure recall timeliness. Complex document structures require vector models to effectively process long texts and integrate multi-source information, distinguishing the importance of different fields. Adverse reaction descriptions, in particular, contain medical terminology and colloquialisms, demanding strong domain adaptability from vector models to accurately understand the relationship between drugs and symptoms, preventing information omission or misjudgment due to vocabulary differences. Precise matching requirements for specific identifiers like drug batch numbers also influence indexing strategy selection.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
embeddingModel | qwen3-embedding-8b or m3e-base | Balances domain adaptability and computational resources, capable of processing medical terminology. |
Chunk size (Segment Length) | 500-800 characters (characters) | Accommodates the detail level of adverse reaction descriptions, balancing context completeness and vectorization efficiency. |
Chunk overlap (Segment Overlap) | 100-150 characters (characters) | Ensures contextual continuity at segment boundaries, improving recall accuracy. |
Recall count (Number of Recalls) | 8-12 entries (items) | Considers multi-source information integration needs, balancing recall breadth and subsequent processing burden. |
Similarity threshold (Similarity Threshold) | Calibrated by actual measurement | Based on DTP pharmacy data characteristics and business needs, to avoid underreporting or false positives, typically between 0.75-0.85. |
Rerank model (Rerank Model) | qwen-max or glm-4 | Improves the relevance of recall results, especially for complex adverse reaction descriptions. |
Common Pitfalls
- After knowledge base construction, retrieval results fail to effectively recall recent adverse reaction events because incremental indexing was not configured or the index update frequency was too low.
- The system fails to match relevant pharmacovigilance information to patient-described adverse reactions using colloquialisms, due to the selected vector model's insufficient understanding of medical terminology and non-standard expressions.
- In FastGPT v4.9.11, configuring the
qwen3-embedding-8bmodel by repeatedly adding models with the same name leads to configuration overwrite, due to not distinguishing between different model configuration instances.
Verification Steps
- Upload a batch of test data containing recent adverse reaction events, perform retrieval, and check if the recall results include the latest information, confirming its timeliness.
- Prepare a set of adverse reaction descriptions containing colloquialisms and medical terminology, perform retrieval, and evaluate the relevance of the recall results to ensure correct understanding of different expression styles.
- Through FastGPT's model management interface, check if the
embeddingModelandreRankModelconfiguration items are correctly loaded and active, ensuring consistent model names and version numbers.
Note: The values provided are common starting points and should be measured against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.