Deployment and Upgrade for Respiratory System Pharmacovigilance

Respiratory system pharmacovigilance data primarily originates from national adverse drug reaction monitoring centers, healthcare institution

Data Characteristics

Respiratory system pharmacovigilance data primarily originates from national adverse drug reaction monitoring centers, healthcare institution reporting systems, pharmaceutical company spontaneous reports, medical literature, and clinical trial data. This data updates frequently, typically weekly or monthly, with significant events triggering immediate updates. The document structure is predominantly unstructured text, including adverse reaction report descriptions, patient medical histories, diagnostic records, and medication details. Some structured data exists, such as drug generic names, dosage forms, dosages, adverse reaction event names, occurrence dates, severity levels, and outcomes. Fields may include standard medical terminology like MedDRA codes and ICD-10 codes. Units involve dosage and frequency representations like mg, ml, and times/day.

Constraints Imposed by These Characteristics on Deployment and Upgrade

The high update frequency and unstructured nature of respiratory system pharmacovigilance data require FastGPT to configure efficient data synchronization mechanisms and robust text processing capabilities during deployment. Frequent data updates mean the knowledge base must support incremental updates and rapidly index new data to avoid duplicate imports. The large volume of unstructured text data, especially involving medical terminology and complex sentence structures, demands more refined segmentation strategies and embedding model choices for FastGPT to ensure accurate extraction and vectorization of core information. Additionally, the presence of specialized codes like MedDRA necessitates considering standardization and entity recognition during data preprocessing to enhance retrieval precision. The deployment environment needs sufficient storage and computing resources to handle continuously growing data volumes and complex vector retrieval tasks.

Configuration Settings

Configuration ItemSuggested ValueRationale
UPLOAD_FILE_MAX_SIZE500 MBRespiratory adverse event reports can contain large amounts of text and attachments, ensuring large files upload successfully.
Chunk size (Segment Length)800 characters (characters)Descriptive text in respiratory adverse event reports is often long; 800 characters helps maintain contextual integrity and reduce information loss.
Chunk Overlap Length (Segment Overlap Length)100 characters (characters)Ensures sufficient overlap between adjacent segments to handle cases where critical information spans across segments, improving retrieval recall.
maxContext8000 tokensRespiratory-related queries often require a longer context to understand complex pathologies and medication situations.
Similarity threshold (Similarity Threshold)0.75For precise matching requirements of medical terms and adverse reaction descriptions, increasing the similarity threshold can reduce false positives.
PARSE_FILE_TIMEOUT_SECONDS600 seconds (seconds)Processing large PDF or DOCX adverse event report documents requires a longer file parsing timeout.

Common Pitfalls

  • After a knowledge base update, user queries still return old data. This occurs because data synchronization tasks did not trigger correctly, or index rebuilding took too long, preventing new data from becoming effective promptly.
  • The system encounters out-of-memory errors when processing certain adverse event reports. This happens when UPLOAD_FILE_MAX_SIZE or PARSE_FILE_TIMEOUT_SECONDS are configured too low, preventing the system from handling oversized files or complex documents.
  • Retrieval recall for certain specific medical terms is unusually low. This is due to an improper segmentation strategy, such as a Chunk size (Segment Length) that is too short, leading to truncated medical entities, or an embedding model that insufficiently understands specialized vocabulary.

How to Verify Correct Configuration

  • Upload a report file containing the latest adverse event data. Query this event via the system interface or API to confirm that new data has been successfully indexed and is retrievable.
  • Simulate a query containing complex medical terminology and multiple adverse reaction symptoms. Check if the recall results include multiple relevant reports and if key entity words (e.g., asthma, bronchitis, cough) are accurately identified.
  • Review system logs for the completion status of file parsing and vectorization tasks, ensuring no timeout or memory-related errors occurred.
  • Use a test set containing common respiratory adverse reactions (e.g., dyspnea, chest tightness). Perform batch retrieval using FastGPT and compare the retrieval results with the expected list of relevant reports to evaluate the recall effect of the Similarity threshold (Similarity Threshold).

The values provided are common starting points and should be measured against specific samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.