Data Characteristics in This Category
Pharmacovigilance (PV) regulatory documents include legal requirements, company Standard Operating Procedures (SOPs), Risk Management Plans (RMPs), Periodic Safety Report (PSR) templates, and training materials. Data sources are diverse, covering guidelines from national drug regulatory authorities, International Council for Harmonisation of Technical Requirements for Pharmaceuticals for Human Use (ICH) documents, and company-specific operational procedures. Document updates depend on regulatory changes, new drug approvals, adverse event report analyses, and internal process optimizations. Core SOPs typically undergo review at least annually, with immediate revisions for regulatory updates or significant events. Document structures are complex, containing specialized terminology, medical abbreviations, charts, and cross-references. Fields and units are highly specialized, such as dosage units (mg, μg, IU), time units (hours, days, weeks), adverse event codes (MedDRA terms), drug batch numbers, and manufacturing dates. Data integrity, accuracy, and traceability requirements are extremely high.
Constraints Imposed by These Characteristics on "Deployment and Upgrade"
The highly specialized and complex nature of pharmacovigilance data imposes specific requirements on FastGPT knowledge base deployment and upgrades. First, diverse document formats (PDF, DOCX, XML, etc.) require FastGPT to have robust document parsing capabilities, ensuring accurate extraction of text content, table information, and chart descriptions. Second, frequent regulatory and SOP updates mean the knowledge base needs to support efficient incremental update mechanisms, avoiding full rebuilds each time and handling document version differences. Third, the extensive specialized terminology and cross-references in documents require a chunking strategy that maintains semantic completeness, preventing critical information from being split. Finally, high demands for data accuracy and traceability mean deployment must focus on index granularity to ensure question-answering results precisely pinpoint original sources. Upgrades must verify lossless migration of knowledge base content.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size | 800–1200 characters | Pharmacovigilance SOP paragraphs are often long, containing multi-step operations or detailed descriptions. This length helps maintain semantic completeness. |
Chunk overlap | 100–200 characters | Ensures contextual continuity between adjacent paragraphs, preventing critical information loss due to splitting, aiding RAG recall. |
maxContext | 4096 | Complex regulatory questions require a longer context window for understanding and reasoning, reducing hallucinations. |
Similarity threshold | 0.75–0.85 | Pharmacovigilance questions demand high answer precision. Increasing the threshold recalls more relevant and accurate knowledge chunks. |
Recall count | 8–12 entries | Regulatory documents may be dispersed across multiple locations. Increasing recall count improves coverage, especially for multi-hop questions. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Parsing large SOP documents takes time. Extending the timeout prevents file loss due to parsing interruptions. |
Three Common Mistakes
- After uploading knowledge base documents, some specialized terms or table contents cannot be accurately recalled in Q&A. This manifests as incomplete answers or discrepancies with the original text. The reason is incorrect document parsing parameter settings, leading to complex structures or specialized terminology not being effectively embedded in the vector database.
- After a system upgrade, the previous Q&A performance degrades, with more "cannot answer" or "outdated information" responses. This may be because during the upgrade process, the index of the old version of the knowledge base was not correctly migrated or was not effectively adapted to the new model version.
- During FastGPT deployment, even with correct external model interface configuration, model loading fails, and the interface displays
Model loading timeout. This usually indicates a network communication issue between the local deployment environment (e.g.,ollama) and the FastGPT container, or the model file is too large, causing loading time to exceed the default limit.
How to Confirm Correct Configuration
- Select a pharmacovigilance SOP document containing complex charts and cross-references. Upload it to the knowledge base and verify that its content is fully parsed, especially whether tables and reference links are recognized.
- Perform an incremental update for recently updated regulations or SOPs. Then, ask relevant questions to confirm that Q&A results reflect the latest content and provide document version information.
- Randomly select 10+ pharmacovigilance questions of varying complexity. Conduct Q&A tests to ensure answer accuracy, completeness, and the ability to pinpoint the original document source. The accuracy rate should exceed a preset threshold.
The values provided are common starting points and should be measured against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.