Data Characteristics
GMP-compliant pharmacovigilance data originates from various sources. These include Adverse Drug Reaction (ADR) reports submitted by Marketing Authorization Holders (MAH), drug quality defect reports, Periodic Safety Update Reports (PSURs), relevant regulatory documents, guidelines, and internal Standard Operating Procedures (SOPs).
Data updates frequently. ADR reports may be added daily, PSURs typically update every six months or annually, and regulatory documents revise periodically. Document formats vary, encompassing structured database records, unstructured PDF documents (e.g., regulatory texts, scanned medical reports), Word documents (e.g., SOPs, investigation reports), and images (e.g., photos of drug packaging defects).
Fields and units are highly specific. Examples include drug names, batch numbers, adverse reaction event terms (using MedDRA coding), report numbers, occurrence dates, dosage units (mg, ml), and administration routes. Medical abbreviations and specific regulatory terms are common.
Constraints on Knowledge Base Retrieval and Recall
The multi-source nature and high update frequency of GMP-compliant pharmacovigilance data require the knowledge base to support efficient data ingestion and incremental updates. This ensures the timeliness of retrieval results.
Image information within unstructured documents, such as images of adverse reaction sites or drug batch number photos, presents challenges for multimodal information processing beyond text. This requires image recognition and text-content association technologies.
The use of specialized terminology, such as MedDRA codes, requires the retrieval model to understand medical semantics. It must support synonym expansion and hierarchical retrieval to prevent recall failures due to terminology mismatches.
The strictness and version control of regulatory documents mandate that retrieval results must trace back to specific regulatory clauses and versions. This imposes stringent requirements on knowledge chunk granularity and metadata management.
The need for precise matching of critical fields like drug dosage and batch numbers limits the applicability of purely semantic retrieval. Keyword or structured queries are necessary in these cases.
Configuration Recommendations
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Chunk Length) | 500–800 characters (characters) | Balances the completeness of regulatory clauses with the information density of a single knowledge chunk, preventing truncation of key information. |
Chunk Overlap Length (Chunk Overlap Length) | 50–100 characters (characters) | Ensures contextual continuity and mitigates potential semantic breaks at chunk boundaries. |
Recall count (Recall Count) | Top 8 entries (top 8) | Considers query complexity and relevance, increasing recall quantity to cover potentially relevant documents, followed by re-ranking optimization. |
Similarity threshold (Similarity Threshold) | Calibrate based on actual measurements | Determined by balancing recall precision and recall rate requirements for specific business scenarios, using an evaluation dataset. |
Rerank result count (Re-ranking Return Count) | Top 3 entries (top 3) | Selects the most relevant knowledge chunks that highly conform to GMP compliance requirements from the recalled results. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds (seconds) | Addresses potentially long processing times for large PDF regulatory documents or scanned document OCR. |
Common Pitfalls
- Model errors or unresponsiveness after knowledge base addition: This manifests as
HTTP 500errors or empty model responses. Potential causes include memory overflow when a locally deployed 14B model loads the knowledge base, or compatibility issues with specific knowledge base component versions (e.g., 4.8.14). - Lack of critical batch number or dosage information in retrieval results: Relevant document content may be semantically related, but precise structured fields are missing. This occurs when the knowledge base processing does not effectively extract and index structured information, or the retrieval model fails to recognize the need for precise matching of specific fields in the query.
- Retrieval results pointing to old versions after regulatory updates: Returned regulatory clauses may be obsolete or revised. This happens when the knowledge base fails to synchronize and process regulatory document version management in a timely manner, lacking effective incremental update mechanisms or metadata version control.
Verification Steps
- Execute retrieval for different query types (e.g., adverse reaction events, specific regulatory clauses, drug batches). Verify the accuracy and completeness of the returned knowledge chunks, confirming the inclusion of key fields such as MedDRA codes, drug names, and regulatory numbers.
- Upload an ADR report containing images and tables. Retrieve the text information within it. Check if the retrieval results correctly identify and recall image descriptions or table content.
- Simulate a regulatory document update. Re-ingest the knowledge base. Then, retrieve pre-update regulatory clauses. Confirm that the system correctly recalls the latest version of the regulatory content and can differentiate between versions.
- Conduct stress tests for high-frequency queries. Monitor knowledge base retrieval service response times and resource utilization. Ensure stable service delivery under high concurrency.
Note: The values provided are common starting points. Measure them against your own samples to determine optimal settings.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.