Data Characteristics in this Category
Rare disease pharmacovigilance data is highly dispersed and heterogeneous. Data sources include reports from global drug regulatory agencies (e.g., FDA Adverse Event Reporting System, FAERS), academic literature, clinical trial data, patient registries, and spontaneous reports from physicians. Update frequencies vary; regulatory databases typically update quarterly or annually in batches, while literature data emerges continuously. Document structures differ significantly, ranging from structured adverse drug reaction reporting forms (MedWatch 3500A) to unstructured clinical notes and case reports. Fields and units are specialized, often including rare disease-specific diagnostic codes (e.g., Orphanet ID), genetic variation information, disease progression scales (e.g., EDSS for MS), and descriptions of extremely low-incidence adverse events. The medical terminology involved is highly specialized, with numerous abbreviations and synonyms.
Constraints Imposed by These Characteristics on "Model Integration and Configuration"
The high dispersion of data requires model integration to include multi-source data interfaces, avoiding the limitations of a single data source. Heterogeneity necessitates more flexible data preprocessing pipelines to accommodate various input formats and structures. Inconsistent update frequencies challenge model real-time capabilities, requiring the configuration of periodic or event-driven data synchronization mechanisms. Diverse document structures demand strong natural language processing capabilities from the model to accurately extract key information from unstructured text, such as adverse events, drug names, and patient characteristics. Rare disease-specific fields and specialized terminology, like Orphanet ID or specific genetic variations, require additional training and vocabulary expansion during knowledge base construction and entity recognition phases. Failure to do so may lead to missed or erroneous information. Identifying extremely low-incidence events requires the model to generalize well to long-tail distribution data.
Configuration Settings
| Configuration Item | Suggested Value | Rationale for this Value |
|---|---|---|
maxContext | 3000 characters | Rare disease case reports are often long, requiring more context to understand relationships. |
Chunk size (Segment Length) | 500–700 characters | Balances semantic completeness with model processing efficiency, preventing truncation of key information. |
Recall count (Recall Count) | 10–15 items | Expands the recall range, increasing the likelihood of discovering rare adverse event patterns. |
Similarity threshold (Similarity Threshold) | 0.75–0.85 | Balances recall and precision, reducing false positives, especially for specialized terminology. |
Rerank result count (Reranked Return Count) | 5 items | Selects the most relevant results for the query, reducing manual review burden. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Handles complex unstructured documents, such as detailed clinical research reports. |
Three Common Pitfalls
- Missing or inaccurate identification of rare disease-specific terms (e.g., specific gene mutation sites) in model output. This occurs because the model's base vocabulary does not adequately cover specialized knowledge in the rare disease domain, or the entity recognition model lacks targeted training.
- After integrating an external knowledge base, the system displays "No available model" or "Model loading failed." This is likely due to an incorrect
MODEL_PROVIDER_API_KEYconfiguration or an invalid format for theEXTERNAL_MODEL_URLof the external model, preventing the platform from making proper calls. - When processing large unstructured documents, a processing timeout or partial content indexing occurs. This usually happens because
PARSE_FILE_TIMEOUT_SECONDSis set too low, not allowing the model sufficient time to complete document parsing and vectorization.
How to Confirm Proper Configuration
- Upload representative rare disease pharmacovigilance reports. Check if the model accurately extracts key entities such as drugs, adverse events, and patient characteristics, and compare the results with human annotations.
- Through the FastGPT model management interface, confirm that the status of integrated external models shows "Available" and that they can successfully execute a simple text generation or embedding task.
- Submit relevant queries for rare adverse reactions or specific patient groups. Verify if the recalled results include relevant information from multiple data sources, and assess whether the recall count and relevance meet expectations.
The values provided are common starting points. Measure them against specific samples to determine optimal settings.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.