Data Characteristics
Small molecule drug pharmacovigilance data originates from clinical trial reports, real-world evidence (RWE), post-market surveillance reports, spontaneous adverse event reporting systems (e.g., FDA FAERS, WHO VigiBase), and medical literature. This data updates frequently; some systems update daily or weekly, while literature data is continuously published. Document structures typically include both structured and unstructured information. Structured data covers patient demographics, drug information (ATC codes, CAS numbers, brand names, generic names), adverse event terms (MedDRA coding), event dates, and outcomes. Unstructured data includes detailed event descriptions, clinical course records, and physician assessments. Fields and units include dosage (often in milligrams (mg), grams (g), or international units (IU)), frequency (e.g., once daily (QD), twice daily (BID)), and time units (days, weeks, months).
Constraints on Deployment and Upgrade
The high update frequency of small molecule drug data requires FastGPT deployments to support incremental updates or periodic full refreshes for data synchronization. This ensures the knowledge base remains current. The mixed document structure (structured and unstructured) necessitates robust text parsing capabilities to extract key entities from free text and effectively embed structured data into retrieval vectors. The presence of specific terminology like MedDRA codes demands that the model understands domain-specific vocabulary and performs accurate entity recognition. The deployment environment needs sufficient storage and computational resources to handle continuous data growth and complex vector embedding calculations. During upgrades, version compatibility is critical. Model updates must not compromise the existing knowledge base's semantic understanding and recall accuracy, especially when processing specific medical terms. Frequent data updates and complex parsing can also lead to lengthy knowledge base index rebuilding times after deployment, affecting service availability.
Configuration Guidelines
| Configuration Item | Suggested Value | Rationale |
|---|---|---|
UPLOAD_FILE_MAX_SIZE | 200 MB | Accommodates large clinical reports and batch imports of literature, preventing upload failures due to file size. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Complex document parsing (e.g., PDF clinical reports) can be time-consuming; prevents parsing timeouts. |
Chunk size | 800–1200 characters | Balances completeness of adverse event descriptions with vector embedding efficiency, preventing semantic fragmentation. |
Recall count | Top 10–15 entries | Ensures enough potentially relevant information is recalled from large datasets, improving subsequent re-ranking accuracy. |
Similarity threshold | 0.75 | Balances recall breadth and precision; too low may introduce noise, too high may miss important information. |
MAX_MEMORY_ALLOCATION | 8 GB | Provides sufficient memory for large knowledge base queries and complex text embeddings, preventing OOM errors. |
Common Pitfalls
- When deleting a large number of documents from a knowledge base folder, the interface displays
timeout of 60000ms exceeded. This occurs because the deletion operation involves extensive index updates, and the default timeout is insufficient, causing the request to be interrupted by the server. - In a local development environment, model loading fails with a
Connection Refusederror. This typically indicates that the model service is not running correctly or there is a port conflict, preventing FastGPT from establishing a connection. - After upgrading to a new version, the relevance of some query results significantly decreases, or critical medical terms are recalled inaccurately. This may be due to changes in how the new model or tokenizer understands domain-specific vocabulary, requiring re-vectorization of the knowledge base.
Verification Steps
- Upload a PDF document containing complex small molecule drug adverse event descriptions. Verify it is parsed and segmented correctly and that its content can be retrieved via keyword search.
- Execute a series of queries containing MedDRA terms. Verify the accuracy and relevance of the recall results, and calibrate the similarity threshold based on business requirements.
- Simulate a large number of concurrent user queries. Monitor system response times and resource utilization to confirm that the deployment environment's stability and performance meet expectations.
- Regularly check knowledge base update task logs. Confirm that the data synchronization mechanism is working correctly and that new data is timely incorporated into the knowledge base and indexed.
Note: The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.