Data Characteristics
Antibody-Drug Conjugate (ADC) pharmacovigilance data originates from clinical trial reports, real-world studies, post-marketing surveillance, and global drug regulatory adverse event databases. This data updates frequently, especially during initial drug launch and when new indications receive approval. Data documents typically include detailed patient demographics, ADC drug dosage and treatment regimens, adverse reaction descriptions (severity, onset time, duration), associated laboratory abnormalities, and causality assessments between adverse reactions and the drug. Fields cover medical terminology (e.g., MedDRA codes), drug batch information, and reporting source identifiers. Units include dosage (mg/kg), time (days, weeks), and biomarker concentrations (ng/mL).
Constraints Imposed by Data Characteristics on Deployment and Upgrade
The multi-source nature and high update frequency of ADC pharmacovigilance data require the deployed FastGPT system to have efficient data ingestion and processing capabilities. Since the data involves sensitive patient information and detailed medical terminology, the system must support strict data anonymization and privacy protection mechanisms. Complex document structures and diverse fields challenge the knowledge base's document parsing and vectorization capabilities, necessitating more refined preprocessing steps to ensure accurate information extraction. The specialized nature of medical terminology demands high foundational medical knowledge and reasoning capabilities from the large language model. Furthermore, regulatory requirements for timely adverse reaction reporting dictate that the upgrade process must support zero-downtime updates or provide rapid rollback mechanisms to ensure service continuity.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
UPLOAD_FILE_MAX_SIZE | 500 MB | Clinical trial reports and post-marketing surveillance documents can contain many images and detailed data, resulting in large file sizes. |
maxContext | 3000 tokens | ADC adverse reaction descriptions are often detailed, requiring a longer context window to capture complete information. |
Chunk size | 800–1200 characters | Ensures each text segment contains sufficient detail for adverse reaction descriptions, facilitating model understanding. |
Recall count | Top 10 entries | Increases the coverage of retrieving relevant adverse reaction cases and medical literature from the knowledge base. |
Similarity threshold | 0.75 | The precision of medical terminology for ADC adverse reactions is high; a lower threshold could introduce irrelevant information. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Processing large PDFs or structured reports can take a long time. |
Common Pitfalls
- Knowledge base query results contain a large amount of irrelevant information because document segment lengths are too short, truncating critical medical descriptions.
- After upgrading to a new version, ONEAPI model connection fails, returning an
HTTP 500error code. This typically occurs when API keys or model endpoint configurations are not correctly migrated in the new version. - The system experiences timeout errors during peak data upload periods, with error logs showing
Connection timed out. This happens when thePARSE_FILE_TIMEOUT_SECONDSparameter is not adjusted for ADC data file sizes.
Verification Steps
- Upload multiple PDF documents containing ADC adverse reaction information. Check if the knowledge base correctly parses and extracts key fields, such as drug name, adverse reaction type, and severity.
- Simulate adverse reaction queries via API or interface. Observe if the returned results include accurate medical terminology and relevant cases. Evaluate if recall and precision meet expectations.
- After a system upgrade, verify all configuration items, especially
ONEAPI_KEYandONEAPI_URL, match the production environment. Confirm model service functionality with a simple API call. - Monitor system logs to ensure no
Connection timed outorHTTP 500errors occur during data ingestion and querying, particularly when processing large files.
Note: The values provided are common starting points. Measure them against your own samples for optimal performance.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.