Data Characteristics for This Category
Rare disease clinical trial pre-screening data primarily originates from global clinical trial registries (e.g., ClinicalTrials.gov, EU Clinical Trials Register), specialized rare disease databases (e.g., Orphanet, OMIM), medical literature (e.g., PubMed, Medline), and pharmaceutical company clinical study reports. Data update frequencies vary; registries may update weekly, while literature databases continuously publish new content. Document structures are highly heterogeneous, including structured trial protocols, patient recruitment criteria, disease diagnostic standards, and unstructured medical reports or patient diaries. Field and unit specificities include disease diagnostic codes (e.g., ICD-10-CM rare disease-specific codes), gene mutation information (e.g., HGVS nomenclature), biomarker concentrations (e.g., ng/mL, pg/mL), and rare disease-specific scale scores (e.g., FARS score, mRS score).
Constraints Imposed by These Characteristics on "Citing Sources and Tracing Origins"
Data source heterogeneity requires FastGPT to support ingestion and parsing of various document formats for source citation and tracing, ensuring all information types are effectively indexed. Inconsistent update frequencies mean knowledge base synchronization mechanisms need flexible configuration to adapt to the timeliness requirements of different data sources. The prevalence of unstructured documents makes deep understanding of text content and key information extraction crucial for accurate tracing. For example, accurately identifying specific gene mutation information from a medical report and linking it to its original source demands higher semantic understanding capabilities from the model. Rare disease-specific fields and units, such as gene sequences or specific scales, must retain their original form and context when cited to avoid information distortion due to insufficient standardization. This directly impacts the model's accuracy in citation generation and tracing paths.
Configuration Settings
| Configuration Item | Suggested Value | Rationale for This Value |
|---|---|---|
Chunk size | 500–800 characters | Rare disease clinical trial documents often contain detailed medical descriptions. A moderate segment length helps maintain contextual integrity without being too long, which could affect recall efficiency. |
Recall count | 8–12 entries | Ensures coverage of sufficient potentially relevant trial information and diagnostic criteria, addressing the issue of insufficient information from a single source. |
Similarity threshold | 0.78–0.85 | Rare disease terminology is highly specialized. A higher threshold filters more precise relevant document segments, reducing noise. |
Rerank result count | 5 entries | After re-ranking, refine to the most relevant few entries, improving final citation quality and user experience. |
ENABLE_NETWORK_SEARCH | true | Supplements the knowledge base with potentially missing latest trial progress or rare disease research dynamics. |
MAX_KNOWLEDGE_BASE_COUNT | 5 | Given the diversity of rare disease data sources, allowing multiple knowledge bases ensures information coverage. |
Three Common Pitfalls
- Symptom: The large language model fails to cite network search results in its answer, even if the network search node appears to be working correctly. Reason: The
ENABLE_NETWORK_SEARCHparameter is not correctly configured totrue, preventing the model from being authorized to use network search results when generating answers. - Symptom: Cited knowledge base document entries do not include specific field values, only listing the document name. Reason: During knowledge base ingestion, key field extraction rules for unstructured documents were not configured or were inaccurate, leading to the loss of important structured information during vectorization and recall.
- Symptom: The model's cited gene mutation information has inconsistent formatting, sometimes missing key loci or mutation types. Reason: During knowledge base import or processing, there is a lack of standardized parsing and storage for rare disease-specific biomedical fields, preventing the model from consistently tracing back to standardized original data.
How to Verify Correct Configuration
- Ask FastGPT questions about specific rare disease clinical trial information. Check if the answer includes citations from multiple knowledge bases and verify that the cited source links are accessible.
- Randomly select several rare disease-related questions. Cross-reference the model's cited content to see if it contains key information such as gene mutations or specific scale scores, and check if the format of this information matches the original document.
- Through FastGPT's management interface, review recent knowledge base synchronization logs to confirm if data updates from different sources are occurring at the expected frequency and without significant errors.
The values given are common starting points and should be measured against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.