Data Characteristics for This Category
Rare disease registration documents draw from diverse data sources. These include clinical trial reports, real-world study data, medical literature, gene sequencing reports, and pharmacokinetic/pharmacodynamic study reports. Data updates are relatively infrequent, typically occurring with new drug development or clinical study results. Document structures are primarily unstructured text, such as PDF clinical study reports and Word expert opinions. Some structured data is also present, like patient registration forms and adverse event reports. Fields and units are highly specialized, for example, gene mutation sites, disease progression scores (e.g., EDSS score), and drug concentrations (unit ng/mL). There are also numerous medical abbreviations and specialized terminology.
Constraints Imposed by These Characteristics on Model Access and Configuration
The unstructured nature of rare disease data requires models with strong text understanding and information extraction capabilities to accurately identify key information from complex documents. Infrequent data updates mean model training and knowledge base construction do not require frequent iteration, but initial construction must ensure historical data coverage. Unique medical abbreviations, specialized terminology, gene sites, and disease scores in documents necessitate special handling during model embedding and entity recognition. This may involve incorporating medical domain pre-trained models or custom vocabularies. The relatively sparse data volume might affect the model's generalization ability in specific rare disease areas, requiring fine-tuning of retrieval strategies and context windows. For numerical data, such as drug concentrations or disease scores, the model needs to understand their meaning and units to avoid misinterpretation.
Configuration Settings
| Configuration Item | Recommended Value | Rationale for Recommendation |
|---|---|---|
maxContext | 8192 | Rare disease documents are often lengthy, requiring a larger context window to capture complete information. |
Chunk size (Segment Length) | 800–1200 characters | Ensures each segment contains sufficient context while preventing excessively long segments that hinder retrieval efficiency. |
Recall count (Number of Retrieved Items) | Top 10–15 items | Given the high information density in rare diseases, increasing the number of retrieved items can improve information coverage. |
Similarity threshold (Similarity Threshold) | Calibrated by actual measurement | Requires adjustment based on the specific dataset to balance recall and accuracy, avoiding interference from irrelevant information. |
Rerank result count (Number of Reranked Items) | Top 5 items | After initial retrieval, a reranking model refines the most relevant content, improving the quality of the final results. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Processing large PDF documents or reports with complex charts requires longer parsing times. |
Three Common Mistakes
- A
422error code during model testing likely indicates that themessagesfield format does not conform to the API specification or contains illegal characters. - Encountering an
[] is too shorterror when starting a conversation usually means the inputmessageslist is empty, indicating the user request did not carry valid conversational content. - Failure to call third-party models like DeepSeek after configuring their API Key might be due to the API Key not being correctly entered in the
API KeyorBase URLfields for the corresponding model in the FastGPT admin backend.
How to Confirm Proper Configuration
- Upload a typical rare disease clinical study report PDF file. Check if the knowledge base successfully parses it and generates text segments.
- Ask questions about specific disease names, gene sites, or drug names from the report. Verify if the model accurately retrieves relevant segments.
- Simulate user queries. Check if the model's answers correctly cite specialized terminology and numerical values from the knowledge base, and confirm unit consistency.
- Test with query statements of varying lengths and complexities. Evaluate the model's response speed and information completeness in different scenarios.
The values provided are common starting points. They should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.