Data Characteristics
Pharmacovigilance data in the infectious disease domain originates from clinical trial reports, real-world studies, spontaneous reporting systems, and medical literature. This data updates frequently; some spontaneous reporting systems may update daily or weekly, while literature data accumulates continuously. Document structures typically include patient demographics, medication history, adverse event descriptions (event name, date, severity, outcome), drug information (generic name, brand name, dosage, administration, batch), and causality assessments. Fields and units require high standardization. For example, adverse events often use MedDRA coding, drug information uses ATC coding, dosage units are precise (e.g., milligrams, international units), and timestamps are precise to the date or even hour.
Constraints Imposed by These Characteristics on Deployment and Upgrade
The high update frequency of infectious disease pharmacovigilance data requires FastGPT deployments to have efficient data ingestion and indexing capabilities. This ensures the knowledge base remains current. The presence of standard coding systems like MedDRA necessitates strict terminology mapping and normalization during data preprocessing. This increases the potential need for PARSE_FILE_TIMEOUT_SECONDS. Complex document structures with many fields require the knowledge base segmentation to effectively preserve relationships between key information, preventing semantic loss. An example is the correspondence between adverse events and implicated drugs. Precise recognition of specific drug dosage units also demands that the underlying text embedding model has a strong understanding of number and unit combinations. This directly influences the selection of embeddingModel and rerankModel.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
UPLOAD_FILE_MAX_SIZE | 500 MB | Accommodates report files containing large amounts of structured data, ensuring smooth file uploads. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Provides sufficient processing time for parsing complex structured documents and performing terminology mapping. |
Chunk size | 800–1200 characters | Balances contextual completeness and retrieval efficiency, ensuring adverse event descriptions and related details are within one segment. |
Recall count | Top 10 entries | Increases recall coverage, capturing various potentially relevant adverse reaction information. |
Similarity threshold | 0.75 | Ensures retrieved results are highly relevant to the query intent, reducing noise. |
rerankModel | bge-reranker-large | Enhances the reranking model's semantic understanding for medical terminology and complex causal relationships. |
Common Pitfalls
- Model channel configuration errors, such as
Invalid API KeyorService Unavailable. This usually indicates incorrect API key configuration for the model provider or network connectivity issues. - Slow knowledge base query responses, even on high-spec servers. This might stem from an inappropriate
embeddingModelorrerankModelselection, or reduced indexing efficiency due to a large volume of knowledge base data. - The Docker-deployed
rerankermodel is inaccessible to FastGPT. This typically occurs when thererankercontainer's network configuration or exposed ports are not correctly mapped, preventing FastGPT from connecting via the specified custom request address.
Verification Steps
- Upload a clinical trial report containing MedDRA codes. Verify that knowledge base segments correctly retain adverse event codes and descriptions.
- Perform a simulated query, such as "arrhythmia caused by azithromycin." Observe if the response time is within expectations and check the relevance of the returned results.
- Review FastGPT's system logs. Confirm no
TimeoutorConnection Refusederrors occurred during data ingestion and querying, especially for model service-related logs.
The values provided are common starting points. Measure them against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.