Data Characteristics for this Category
Gene therapy AAV (adeno-associated virus vector) pharmacovigilance data originates from clinical trial reports, real-world studies, post-market surveillance, and global drug regulatory databases (e.g., FDA, EMA). Data updates are frequent. During clinical trials, safety reports may be submitted weekly or monthly. Post-market, data may be summarized quarterly or annually based on event severity and frequency. Document structures are complex. They typically include structured adverse event reports (e.g., MedDRA coding), unstructured clinical narratives, patient histories, laboratory results, imaging data, and treatment details. Beyond standard drug and patient information, fields and units also cover specific details like AAV vector serotype, dosage, administration route, target cell type, gene expression product, immunogenicity assessment indicators (e.g., anti-AAV antibody titers), and genomic integration site analysis. Measurement units include viral genomes (vg) per kilogram of body weight and titer units.
Constraints from these Characteristics on "Deployment and Upgrade"
High-frequency data updates require FastGPT's data synchronization mechanism to be highly real-time. This ensures timely capture of alert information and risk signals. Complex document structures and multi-source heterogeneous data (structured and unstructured coexistence) challenge data preprocessing and knowledge extraction. This demands robust text parsing capabilities and flexible knowledge base construction strategies. The presence of specific fields, such as AAV vector serotype or immunogenicity indicators, means that knowledge segmentation and vectorization require specialized entity recognition and semantic understanding models. This prevents loss of critical information or semantic bias. Additionally, the ability to parse regulatory agency-specific report formats (e.g., E2B reports) is crucial for deployment. The deployment environment must support high concurrency and elastic scaling to handle sudden data processing demands. The upgrade process needs to ensure model compatibility and smooth data migration, reducing business interruption risks.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Processing PDF or DOCX files containing extensive clinical narratives and complex tables requires a longer parsing time. |
Chunk size | 800 characters | Ensures complete MedDRA terms, immunological indicators, and relevant context are included, while avoiding information redundancy. |
Recall count | 10 entries | Increases the probability of recalling all relevant risk signals from a large volume of adverse event reports. |
Similarity threshold | 0.78 | Distinguishes subtle clinical presentation differences and identifies potential rare or atypical adverse reactions. |
Rerank result count | 5 entries | Focuses on the most relevant key information through re-ranking models, building on a high recall rate. |
UPLOAD_FILE_MAX_SIZE | 500 MB | Supports uploading large clinical study reports and case files, including multimedia content. |
Three Common Mistakes
- After data ingestion, the interface displays in English: The
FASTGPT_LOCALEenvironment variable in thedocker-compose.ymlfile is usually not correctly set tozh-CN, or the container restart did not apply the change. - Voice input function is unusable: Even if the Whisper model is deployed, FastGPT's
docker-compose.ymlneeds to configure the correctTTS_URLorASR_URLpointing to the Whisper service address, and network accessibility must be ensured. - Login errors after re-pulling images: This may be due to incompatibility between the old data structure and the new version. Check if database migration scripts ran successfully, or back up and clear old data before upgrading.
How to Confirm Correct Configuration
- Upload a clinical report PDF containing AAV vector serotype, dosage, and immunogenicity indicators. Check if it is correctly parsed and if key fields are extracted.
- Query a specific adverse event (e.g., "thrombocytopenia") via the FastGPT interface. Verify if the returned results include information on relevant AAV drug dosage, administration route, and patient immune status.
- Simulate a new adverse event report data import. Check the real-time nature of knowledge base updates and confirm if new risk signals are identified and presented promptly.
- Configure a voice input query in FastGPT, such as asking "association between AAV8 vector and hepatotoxicity." Verify the accuracy of speech recognition and knowledge recall.
The values provided are common starting points and should be measured against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.