HTTP API and External Systems for Rare Disease Pharmacovigilance

Rare disease pharmacovigilance data has distinct characteristics. Data sources are highly dispersed, typically including national or regional rare

Data Characteristics in this Domain

Rare disease pharmacovigilance data has distinct characteristics. Data sources are highly dispersed, typically including national or regional rare disease registries, clinical trial reports, real-world evidence (RWE) studies, medical literature case reports, and patient reports. Due to the small patient population, the incidence of adverse events (AEs) or adverse drug reactions (ADRs) is extremely low. Data update frequency is less intensive than for common disease medications, potentially occurring quarterly or annually. Document structures are diverse, encompassing both structured database records and a large volume of unstructured free text, such as physician consultation notes and patient self-reports. In terms of fields, in addition to standard patient ID, drug name, dosage, route of administration, adverse event description, onset time, and severity, data often includes genetic information, genetic test results, diagnostic criteria (e.g., specific gene mutations or enzyme activity levels), and family history. Units must strictly adhere to medical standards, for example, dosage units (mg/kg) and time units (days, weeks, months).

Constraints Imposed by these Characteristics on "HTTP API and External Systems"

The high dispersion of rare disease data requires HTTP API designs to aggregate data from multiple heterogeneous external systems. Due to lower update frequencies, API calling strategies need adjustment to avoid unnecessary frequent polling; event-driven or scheduled batch synchronization can be employed. The large amount of unstructured text demands higher capabilities for data preprocessing and knowledge extraction, requiring external systems with robust natural language processing (NLP) capabilities. Special fields like genetic information require APIs to support complex data types or nested structures during data transmission, ensuring data integrity and accuracy. Critical information such as diagnostic criteria may exist as text descriptions, necessitating API calls to external medical knowledge graphs or ontology services for standardization and encoding. Furthermore, because data volume is relatively small but value density is high, strict requirements are placed on API stability and error handling mechanisms to prevent data loss or misreporting.

Configuration Settings

Configuration ItemRecommended ValueRationale
maxContext2048 charactersRare disease case reports often contain detailed medical history and diagnostic criteria, requiring sufficient context length to understand complex medical backgrounds.
chunkOverlapRatio0.15Considering the coherence of medical terminology and narratives, appropriate overlap helps maintain the integrity of key information, preventing semantic breaks due to splitting.
PARSE_FILE_TIMEOUT_SECONDS300 secondsProcessing files containing large amounts of unstructured text (e.g., clinical reports in PDF format) requires longer parsing times to prevent file processing failures due to timeouts.
similarityThreshold0.05Recalling rare disease-related information requires higher sensitivity; even lower relevance may contain critical clues.
external_api_timeout60 secondsWhen calling external genetic test report parsing services or medical ontology mapping services, these services can be computationally intensive, requiring ample response time.
MAX_RETRIES3 timesExternal systems (e.g., national rare disease registries) may experience intermittent network fluctuations; increasing retry attempts improves data synchronization success rates.

Three Common Mistakes

  • When calling an external API, a 400 or 500 status code is returned without detailed error information, making it difficult to determine if it's a parameter format error or an internal error of the external service.
  • Knowledge base data synchronized from an external business system fails to return the expected referenced knowledge base ID when retrieved in FastGPT. This occurs because the original knowledge base ID was not mapped to FastGPT's datasetId field during the API call.
  • After uploading PDF documents containing many medical terms and rare disease names, text segmentation results are unsatisfactory. This is because the tokenizer is not optimized for domain-specific vocabulary, leading to incorrect splitting of key terms.

How to Verify the Setup

  • Submit a knowledge document containing a rare disease case to FastGPT via HTTP API. Observe if the returned status code is 200 and check if FastGPT's backend successfully created or updated the corresponding knowledge entry.
  • Use FastGPT's chat interface to query questions related to the newly submitted knowledge document. Verify that the returned results include the correct knowledge base ID and cited text snippets, and cross-reference them with the original document content.
  • Simulate an external system call, transmitting data that includes genetic information and complex medical terminology. Check FastGPT's receiving logs to confirm that all key fields are correctly parsed and stored, without data truncation or format errors.

The values provided are common starting points and should be measured against the reader's own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.