Data Characteristics in this Category
Pharmacovigilance (PV) submission documents contain extensive structured and unstructured data. Data sources are diverse, including clinical trial reports, real-world evidence (RWE), individual case safety reports (ICSRs), literature reviews, drug inserts, risk management plans (RMP), and regulatory guidelines. Data update frequencies vary. ICSRs may update daily or even in real-time, while clinical study reports or RMPs might update quarterly or annually. Document types include PDF, Word, Excel, and XML. Fields are highly specific. For instance, ICSRs include MedDRA codes and WHO-DD codes. Dosage units involve mg, μg/kg, and time units are precise to hours and minutes, often involving the time difference between event occurrence and reporting.
Constraints Imposed by these Characteristics on Reference and Traceability
The diversity and update frequency of pharmacovigilance data impose specific requirements on reference and traceability. First, multi-source heterogeneous document formats require robust file parsing capabilities to ensure complete information extraction, especially for tables and image text within PDFs. Second, high-frequency updates for data like ICSRs necessitate that the RAG system quickly indexes and maintains an up-to-date knowledge base, preventing the citation of outdated information. The presence of specialized coding systems like MedDRA and WHO-DD makes precise retrieval and matching critical. Standard keyword matching may lead to inaccurate recall. Furthermore, for numerical fields with units, such as dosage and time, the RAG system must understand context to differentiate numerical values under different units. This prevents miscitation or incorrect traceability. For example, a document with a dose of "10 mg" and another with "10 μg/kg" may appear numerically similar but have vastly different actual meanings.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size | 800–1200 characters | Accommodates long paragraph descriptions and complex logic in pharmacovigilance documents, ensuring contextual completeness. |
Chunk overlap | 200 characters | Ensures critical information across paragraphs is not lost, especially in adverse event descriptions. |
Recall count | Top 5 entries | Balances retrieval efficiency and coverage, prioritizing the retrieval of a few highly relevant, high-quality documents. |
Similarity threshold | 0.75 | Pharmacovigilance demands high accuracy in citations. A threshold that is too low introduces noise, while one that is too high may lead to omissions. |
Rerank result count | Top 3 entries | Further refines recall results, improving the quality of citations presented to the user. |
maxContext | 4000 token | Ensures the capacity to hold detailed event descriptions and background information found in pharmacovigilance reports. |
Three Common Pitfalls
- After entering a Chinese query, English documents in the knowledge base are not cited. This occurs when the vector model or retriever fails to effectively handle cross-language semantic matching.
- The query content has low relevance to the knowledge base content, yet the system returns irrelevant information. This likely happens when the
Similarity threshold(similarity threshold) is set too low, leading to the recall of significant noise. - When parsing complex PDF documents, some table data or text within images are not recognized, resulting in missing information. This usually relates to insufficient support for specific formats by the file parser or a
PARSE_FILE_TIMEOUT_SECONDSsetting that is too short.
How to Confirm Proper Configuration
- Select typical pharmacovigilance queries and verify whether the cited documents contain the query keywords and relevant professional terminology.
- For queries including MedDRA codes and WHO-DD codes, check if the system's returned citation sources accurately point to the original text containing these codes.
- Evaluate the system's citation of numerical fields (e.g., dosage, time) to confirm that units and values match the original text without misinterpretation or confusion.
- After a knowledge base update, check if frequently changing data (e.g., the latest ICSRs) are promptly indexed and used for subsequent citations.
The values provided are common starting points. Measure them against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.