Data Characteristics
Pharmacovigilance data originates from post-market adverse event reports (ADR/SAE), safety updates from global regulatory agencies, medical literature, clinical trial data, and internal risk management plans. This data updates frequently. Adverse event reports may have daily additions. Regulatory agency updates typically occur weekly or monthly. Document structures include Individual Case Safety Reports (ICSRs), aggregate reports (e.g., PSURs, DSURs), Risk Evaluation and Mitigation Strategies (REMS/RMP), and Standard Operating Procedures (SOPs). ICSRs usually contain structured fields like patient demographics, drug information, adverse event details, and outcomes, along with medical narrative text. Aggregate reports include extensive statistical data, signal detection results, and professional medical analyses. Field units include dosage units (mg, IU), frequency units (times/day), and time units (days, weeks, months).
Constraints on Citation and Traceability
The high update frequency of pharmacovigilance data requires rapid synchronization and indexing capabilities for the knowledge base to ensure timely citation. Documents containing both structured data and unstructured text necessitate a combination of keyword matching and semantic understanding for retrieval. The concise narrative text in ICSRs and the longer medical analyses in aggregate reports demand varied text segmentation strategies. Overly long segments may dilute key information, while overly short segments may lose context. Data sensitivity and regulatory compliance mandate that all citations must be accurate and traceable to original documents, preventing misinterpretation or alteration of information. Traceability is critical, especially when responding to regulatory inquiries. Differences in structured field units also require correct identification and presentation during information extraction.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
maxContext | 800–1200 characters | Balances the completeness of adverse event narratives with the conciseness of aggregate reports |
Similarity threshold (Similarity Threshold) | 0.75 | Ensures high relevance of retrieval results to pharmacovigilance queries, reducing noise |
Recall count (Recall Count) | Top 5 items | Covers key information points while avoiding excessive redundancy |
Rerank result count (Reranked Return Count) | 3 items | Prioritizes the most relevant and traceable short sentences or paragraphs |
Chunk size (Segment Length) | 300 characters | Suitable for extracting short texts from ICSRs and key paragraphs from aggregate reports |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Addresses parsing time for large aggregate reports or batch processing of multiple documents |
Common Pitfalls
- Replies contain only citation information without generating a specific answer. This usually results from overly conservative generation model parameters or unclear generation requirements in the prompt.
- Knowledge base citation variables cannot be selected in code execution nodes. This occurs when knowledge base query results are not correctly bound to variables, or the variable type does not match expectations.
- Citation templates fail to effectively integrate multi-turn dialogue context, leading to disjointed answers. This may relate to incorrect use of historical dialogue variables or context window limitations in the template.
Verification Steps
- Submit typical pharmacovigilance queries. Check if the generated answer includes knowledge base citation links and if clicking the links correctly navigates to the original documents.
- Test with a pharmacovigilance report containing both structured data and long narratives. Verify if the specific content, units, and fields cited in the answer match the original text.
- Simulate regulatory inquiry scenarios. Ask questions about specific adverse events or drug safety signals. Evaluate the accuracy, completeness, and ease of citation traceability in the answers.
- After a knowledge base update, perform the same queries again. Verify if the citation information reflects the latest data and check the efficiency of updated document parsing and indexing.
The values provided are common starting points and should be measured against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.