Data Characteristics in this Category
Supplier audit data in pharmacovigilance primarily consists of audit reports, supplier qualification documents, and adverse drug reaction (ADR) reports. Audit reports are typically in PDF or Word format. They contain audit findings, corrective and preventive actions (CAPA), and completion status. Update frequency aligns with the audit cycle, usually annually or biennially. Supplier qualification documents, such as manufacturing licenses and quality management system certifications (e.g., ISO 9001, GMP certificates), are often scanned images or electronic documents. Their update frequency depends on certificate validity, typically 3-5 years. ADR reports come from various sources, including clinical trial data, post-market surveillance data, and spontaneous reporting systems. Their data structure varies, potentially including patient demographics, drug information, adverse event descriptions (often free text), severity ratings, and causality assessments. ADR data has a high update frequency, with new entries possibly daily.
Constraints on Tool Calling and Plugins from these Characteristics
The heterogeneity and varied update frequency of supplier audit data directly impact tool calling and plugin design. The PDF/Word format of audit reports and qualification documents requires file parsing tools to effectively extract structured information. Accurate recognition of tables and nested text is crucial, potentially necessitating specialized OCR or document parsing plugins. Due to infrequent updates, the indexing strategy for this data in the knowledge base can involve periodic full or incremental updates. The free-text descriptions in ADR reports demand advanced natural language processing (NLP) capabilities. Plugins need to handle medical terminology, abbreviations, and non-standard expressions, extracting key entities and relationships. High-frequency ADR data updates require tool calling to support real-time or near real-time data ingestion, rapidly triggering incremental indexing and retrieval in the knowledge base. Additionally, inconsistent fields and units across different data sources require standardization or mapping at the tool calling layer.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale for this Value |
|---|---|---|
Chunk size (Segment Length) | 500–800 characters | Ensures each segment contains sufficient context, preventing semantic loss from excessive splitting, and accommodates the paragraph length of audit reports. |
Overlap Length | 50–100 characters | Guarantees contextual continuity between segments, especially when processing lengthy audit findings or adverse drug event descriptions. |
Max File Size | 100 MB | Accommodates the upload requirements of large audit reports and scanned qualification documents, preventing upload failures due to oversized files. |
Parse Timeout | 300 seconds | Allows ample time for parsing PDF files that contain numerous images or complex tables. |
Recall count (Recall Count) | Top 8 entries | Pharmacovigilance reviews typically require synthesizing information from multiple sources; increasing recall count can improve result comprehensiveness. |
Similarity threshold (Similarity Threshold) | 0.75 | Balances recall and precision, ensuring retrieved results are highly relevant to the query and avoiding interference from irrelevant information. |
Three Common Mistakes
- Tool calling returns a
404error: This typically occurs due to incorrect backend API addresses or routes configured in OneAPI, or a mismatch between the URL called in the FastGPT plugin and the actual service address. - Plugin execution returns empty or incomplete results: This might happen if the document parsing plugin fails to correctly identify table data in PDFs or scanned documents, leading to critical field extraction failures.
- RAG results do not include the latest ADR information: The knowledge base update strategy is not synchronized with the ADR data source's update frequency, causing indexed data to lag behind actual adverse events.
How to Confirm Correct Setup
- Upload typical audit reports (PDF/Word) and qualification document samples. Check if the knowledge base correctly segments and extracts key information, such as audit findings and CAPA status.
- Design queries containing medical terminology and adverse event descriptions. Verify that tool calling correctly identifies and invokes relevant plugins, and recalls relevant and complete reports from the ADR knowledge base.
- Simulate new ADR data ingestion and query shortly thereafter. Confirm that the knowledge base's incremental update mechanism works as expected and that new data is retrievable.
Note: The values provided are common starting points. Measure against your own samples for optimal configuration.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.