Data Characteristics for this Category
Solid tumor pharmacovigilance data typically originates from clinical trial reports, real-world studies, adverse event reporting systems (e.g., FAERS, EudraVigilance), and medical literature. Data update frequencies vary; clinical trial data is usually released centrally after studies conclude, while adverse event reports are continuous, potentially updating daily. Document structures are diverse, including structured database records, unstructured free-text descriptions (e.g., case reports, physician notes), and semi-structured PDF reports. Beyond general patient and drug information, specific fields include tumor type (e.g., ICD-O-3 codes), tumor stage, comorbidities, and tumor treatment history. Adverse reaction descriptions often involve medical terminology, abbreviations, and may include units and quantitative information such as dosage, administration route, and reaction severity (e.g., CTCAE grades).
Constraints Imposed by These Characteristics on "Deployment and Upgrades"
The multi-source and heterogeneous nature of solid tumor pharmacovigilance data demands robust data ingestion and preprocessing capabilities during deployment. A high proportion of unstructured data necessitates powerful text parsing and extraction, requiring sufficient file processing resources during deployment. The continuous data update characteristic requires systems to support incremental learning and dynamic knowledge base updates. Upgrade processes must consider zero-downtime updates and data consistency maintenance. Specific fields like tumor type and staging require customized entity recognition models, impacting memory requirements for model loading and inference. The use of medical terminology and abbreviations challenges the domain adaptability of embedding and retrieval models, potentially requiring domain-specific pre-trained models or fine-tuning. Furthermore, quantitative information on adverse reaction severity requires models to accurately understand and utilize numerical values with units, which needs special attention in knowledge base chunking and retrieval strategies.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
UPLOAD_FILE_MAX_SIZE | 500 MB | Handles clinical report PDFs containing large images or scanned documents |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Accommodates parsing time for complex structures or extremely long case reports |
Chunk Length | 800–1200 characters | Balances contextual completeness for solid tumor reports with retrieval efficiency |
maxContext | 32000 | Covers contextual information from lengthy medical literature and multiple adverse event reports |
Similarity Threshold | 0.75 | Ensures high recall, capturing potentially relevant adverse reaction information |
Rerank Return Count | Top 8 | Reduces processing burden on subsequent models while maintaining relevance |
Three Common Pitfalls
- After local Docker Compose deployment, OneAPI cannot access FastGPT interfaces, and logs show
Invaild url. A common cause is incorrect configuration of theFASTGPT_URLenvironment variable, preventing OneAPI from resolving FastGPT's internal service address. - Uploading large PDF files results in an upload failure or prolonged unresponsiveness in the interface. This typically occurs because the
PARSE_FILE_TIMEOUT_SECONDSparameter is set too low, leading to file parsing timeouts, orUPLOAD_FILE_MAX_SIZElimits the file size. - Knowledge base Q&A retrieval quality is poor, or irrelevant information is returned. This may be due to an excessively long
Chunk Lengthcausing context loss, or aSimilarity Thresholdset too low, introducing noisy data.
How to Verify Correct Configuration
- Upload and parse a solid tumor clinical report PDF containing key fields like tumor stage and
CTCAEgrade. Verify that the knowledge base correctly extracts and stores this information. - Perform a knowledge base Q&A query against a document known to contain a specific adverse reaction (e.g.,
immune-related pneumonitis,CTCAEgrade3). Confirm that the returned results accurately mention the adverse reaction and its severity. - Simulate high-concurrency scenarios for knowledge base retrieval and Q&A. Observe if system response times meet expectations and check logs for file parsing timeouts or out-of-memory errors.
- After a FastGPT upgrade, compare Q&A results for the same set of solid tumor-related queries before and after the upgrade to assess the stability or improvement of the knowledge base and model performance.
Note: The values provided are common starting points. It is important to measure and adjust them based on specific data samples and operational requirements.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.