Data Characteristics
Pharmacovigilance data for small molecule drugs originates from clinical trial reports, real-world evidence (RWE), post-market surveillance, medical literature, and spontaneous patient reports. This data updates frequently, especially during a drug's initial market release, with new adverse event reports potentially generated weekly or even daily. Document structures typically include various types: structured reporting forms (e.g., ICH E2B format), semi-structured medical texts (e.g., patient records, diagnostic reports), and unstructured free-text descriptions. Key fields include drug name, active ingredient, dosage form, dose, administration route, adverse event (MedDRA coding), onset time, severity, outcome, relevant medical history, and concomitant medications. Units commonly used are milligrams (mg), grams (g), milliliters (mL) for dosage, and hours, days, weeks, months, years for time.
Constraints Imposed by Data Characteristics on Vector Models and Indexing
The high frequency of updates in small molecule drug adverse event reports demands that vector indexes possess efficient incremental update capabilities to ensure timely retrieval results. Data diversity, particularly the mix of structured data and free text, requires vector models to process and effectively integrate both types of information. The use of specialized terminology like MedDRA codes places higher demands on the model's semantic understanding to accurately capture medical concepts. Furthermore, adverse event descriptions often include time series and causal relationships, requiring the vectorization process to capture inter-event associations. Accurate matching and recall of key fields such as drug name, dose, and adverse events are fundamental to effective pharmacovigilance, making precise vector recall critical.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Chunk Length) | 300–500 characters (characters) | Accommodates the typical length of factual descriptions in adverse event reports, avoiding information fragmentation. |
Chunk Overlap Length (Overlap Length) | 50–80 characters (characters) | Ensures context continuity and improves the accuracy of semantic connections across chunks. |
Recall count (Recall Count) | Top 10–20 entries (top 10–20 items) | Balances recall rate with computational resource consumption, covering potentially relevant information. |
Similarity threshold (Similarity Threshold) | Calibrated by empirical measurement | Requires evaluation using F1 score based on specific datasets and business needs. |
embeddingModel | bce-embedding | Provides good semantic understanding for Chinese medical texts. |
Index Update Strategy | incremental update, daily batch processing | Addresses high-frequency data updates, ensuring data timeliness. |
Common Pitfalls
- Knowledge base search takes too long, manifesting as query requests with prolonged unresponsiveness or timeouts. This may be due to the chosen vector model having high computational resource requirements, while hardware configurations (e.g., CPU or GPU memory) are insufficient.
- Retrieval results contain a large number of irrelevant or low-relevance documents. This may be due to a
Similarity threshold(similarity threshold) set too low, leading to an overly broad recall scope. - Despite the upstream embedding service running normally, FastGPT reports "No Available channel" (no available channel). This usually means that the channel pointed to by
ONEAPI_URLorFASTGPT_EMBEDDING_MODELconfiguration items is not correctly mapped to an available embedding service.
Verification Steps
- Using the FastGPT management interface, select the configured
embeddingModeland perform vectorization tests on typical adverse event reports. Observe if theembeddingfield generates a reasonable numerical array. - Upload a batch of documents containing known adverse events. Use relevant drug names or adverse event symptoms to perform retrieval and check if the expected documents are included in the recall results, evaluating their ranking.
- Monitor FastGPT operation logs to confirm that no connection errors or authentication failures related to the
embeddingservice occur during knowledge base updates or queries. - For critical adverse event queries, perform multiple retrievals and adjust the
Similarity threshold(similarity threshold) based on feedback from business experts regarding the relevance of retrieval results until satisfactory.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.