Data Characteristics
Cardiovascular pharmacovigilance data originates from clinical trial reports, real-world studies, adverse event reporting systems (e.g., FDA FAERS, EMA EudraVigilance), medical literature, drug labels, and professional guidelines. This data updates frequently; adverse event reporting systems may update daily, while literature and guidelines update quarterly or annually. Document structures vary, including structured database records, semi-structured XML or JSON files, and unstructured text reports (e.g., case narratives). Key fields include drug generic name, brand name, indication, adverse event name, occurrence time, severity, patient basic information (age, gender, concomitant medications), and event description. Units include milligrams (mg), grams (g), or international units (IU) for dosage, and days, weeks, months, or years for time.
Constraints on Knowledge Base Retrieval and Recall
The high update frequency of cardiovascular pharmacovigilance data requires an efficient incremental update mechanism for the knowledge base to ensure timely retrieval results. Diverse and heterogeneous data structures necessitate flexible data preprocessing and vectorization strategies for unified representation and effective retrieval. Unstructured text reports often contain colloquialisms mixed with medical terminology in adverse event descriptions, posing challenges for semantic understanding and entity recognition. The complexity of cardiovascular diseases often involves multiple co-administered drugs, and adverse events may relate to multiple factors, requiring the retrieval system to identify multi-entity relationships and support complex queries. Patient individual differences are significant, so retrieval results must consider contextual information like age and underlying conditions to improve recall relevance.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Chunk Size) | 500–800 characters | Balances semantic completeness and vectorization efficiency. Avoids diluting key information with overly long texts or losing context with overly short texts. |
Chunk Overlap Length (Chunk Overlap) | 50–100 characters | Ensures contextual continuity at chunk boundaries, reducing the risk of critical information being split, especially for adverse event descriptions. |
Recall count (Recall Count) | top 10–15 items | Balances recall breadth with the processing load of subsequent reranking models, ensuring coverage of potentially relevant information. |
Similarity threshold (Similarity Threshold) | 0.75–0.85 | Increases the relevance of recall results and reduces noise, specifically for cardiovascular pharmacovigilance data. |
Rerank model (Reranker Model) | BGE-reranker-large | Improves the ranking accuracy of retrieved documents, particularly for identifying the most relevant medical reports from a large initial recall set. |
Rerank result count (Reranked Return Count) | top 5 items | Focuses on the few most relevant documents after reranking for subsequent large language model processing, improving efficiency and accuracy. |
Common Pitfalls
- Insufficient knowledge base retrieval results, failing to cover all relevant adverse event reports. This may occur if the
Similarity threshold(Similarity Threshold) is set too high orRecall count(Recall Count) is too low, causing the system to prematurely filter out potentially related documents. - Retrieved adverse event descriptions do not match the actual query intent or contain excessive irrelevant information. This usually results from an inappropriate
Chunk size(Chunk Size) setting, failing to capture core semantics effectively, or not fully utilizing the reranking model to optimize ranking. - Locally deployed
rerankermodels, such as Ollama, fail to load or run, displaying error messages indicating missing model files or configuration errors. This is due to incorrect model path specification or uninstalled dependencies.
Verification Steps
- For typical queries, check if the raw recall count from the knowledge base covers the expected relevant documents and observe the distribution of relevance between recalled documents and the query.
- Simulate queries for different severities and types of adverse drug reactions. Evaluate the reranked result list and confirm that the most relevant documents appear at the top.
- Regularly test with recently published cardiovascular pharmacovigilance data to verify if the knowledge base's update mechanism reflects the latest information promptly and ensures the timeliness of retrieval results.
Note: The values provided are common starting points. Measure performance against your own samples to determine optimal settings.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.