Data Characteristics in This Domain
Clinical Decision Support (CDS) in pharmacovigilance primarily uses data from drug labels, drug interaction databases, adverse event reports (e.g., FDA Adverse Event Reporting System, FAERS), clinical guidelines, pharmacology literature, and professional journals. Update frequencies vary. Drug labels and clinical guidelines typically update quarterly or annually, while adverse event reports can be real-time. Document structures for drug labels include standard sections like ingredients, indications, contraindications, dosage, administration, and adverse reactions. Drug interaction databases primarily use structured data, describing drug-drug, drug-food, and drug-disease interaction types and severity. Fields and units include drug names, dosages (milligrams, milliliters), administration routes, frequencies, adverse event terms (using MedDRA codes), patient age, and gender.
Constraints on Knowledge Base Retrieval and Recall
The heterogeneous nature of pharmacovigilance data requires knowledge base retrieval to handle diverse structures and update frequencies. Long text in drug labels and clinical guidelines necessitates effective knowledge chunking strategies to capture critical information and prevent loss of important context. For example, adverse reaction sections may contain extensive descriptive text, requiring fine-grained chunking for precise retrieval. The structured nature of drug interaction databases demands advanced semantic understanding from the retrieval system to identify drug names and interaction types. Varying update frequencies mean the knowledge base must support incremental updates and version management to ensure timely retrieval results. Standardization of fields and units, such as MedDRA codes, requires the retrieval system to understand and match these specialized terms, preventing retrieval failures due to synonyms or near-synonyms. Long-tail rare adverse events challenge the robustness of recall algorithms, requiring stronger generalization capabilities.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Chunk Length) | 300–500 characters (characters) | Balances the completeness of adverse reaction descriptions with the information density of a single chunk, preventing overly long chunks from diluting the topic. |
Chunk Overlap Length (Chunk Overlap Length) | 50–100 characters (characters) | Ensures contextual continuity, especially when critical information like adverse reactions or contraindications spans across chunks. |
Recall count (Recall Count) | 5–8 entries (items) | Considers the complexity of pharmacovigilance decisions, requiring multi-faceted information support, but too many items increase model burden. |
Similarity threshold (Similarity Threshold) | Calibrate by actual measurement (Calibrated by actual measurement) | Requires iterative testing with real query cases to balance recall and precision, ensuring rare adverse reactions are also recalled. |
Rerank result count (Rerank Return Count) | 3 entries (items) | Further refines the most relevant items from the initial recall, reducing noise. |
Enable Semantic Chunking | Yes | Better identifies the logical structure and paragraph boundaries in documents like drug labels and clinical guidelines, improving chunking quality. |
Three Common Pitfalls
- Retrieval results contain a large amount of irrelevant information. This typically occurs when the
Similarity threshold(Similarity Threshold) is set too low, leading to the recall of non-strongly related content. - Queries for a specific adverse reaction of a particular drug do not include detailed information about that adverse reaction. This may be because knowledge chunking did not adequately preserve the complete description of the adverse reaction, or the
Chunk size(Chunk Length) was set too short, truncating critical information. - API calls sometimes return empty results or generic responses, even when relevant content exists in the knowledge base. This could be due to API request parameters (e.g., the
queryfield) not matching the knowledge base content, or a lowRecall count(Recall Count) filtering out relevant but less similar information.
How to Verify Configuration
- Select representative pharmacovigilance queries to simulate real-world questions. Check if the recall results include all relevant drug labels, interaction information, and adverse event reports. Verify that key fields (e.g., drug name, adverse reaction MedDRA code) match accurately.
- For known drug adverse event cases, design queries from multiple angles. Observe the consistency and completeness of recall results across different query formulations.
- Use the FastGPT backend debugging tools to view the
Recall count(Recall Count),Chunk Content, andSimilarity Scorefor each query. Analyze which chunks are recalled and their relevance. Adjust theSimilarity threshold(Similarity Threshold) based on actual business needs. - Regularly import new drug labels or adverse event data. Perform query tests to verify the timeliness and accuracy of retrieval after knowledge base updates.
Note: The values provided are common starting points. Measure against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.