Data Characteristics for This Category
Rare disease product data primarily originates from clinical trial reports, drug monographs, medical literature, patient registries, and information released by government regulatory bodies. This data updates relatively infrequently, typically quarterly or annually, coinciding with new drug approvals, expanded indications, or significant research findings. Document structures are complex, containing extensive unstructured text such as disease mechanism descriptions, clinical symptoms, diagnostic criteria, treatment protocols, drug interactions, and adverse reactions. Structured data includes gene loci, protein expression, drug dosages, and treatment cycles. Fields often contain specialized medical terminology and International Classification of Diseases (ICD) codes. Units involve mg/kg, IU/mL, mM, etc., demanding high precision and standardization.
Constraints Imposed by These Characteristics on "Model Integration and Configuration"
The unstructured nature of rare disease data requires models with strong text understanding and information extraction capabilities to accurately identify key entities and relationships from complex medical texts. The low data update frequency means model training and fine-tuning cycles can be relatively flexible, but knowledge base maintenance needs to focus on the accuracy and integration of new data. The presence of specialized terminology and coding systems necessitates dedicated preprocessing and domain knowledge injection for vocabulary and entity recognition. Data sensitivity demands strict adherence to data privacy and security protocols during model deployment, especially in scenarios involving patient information. High-precision unit and dosage information places strict requirements on model output accuracy, requiring consideration of numerical validation and unit conversion during configuration.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
maxContext | 8192 | Rare disease documents are often long, requiring a larger context window to capture complete information. |
Chunk size (Chunk Size) | 1000 characters (characters) | Balances semantic completeness with retrieval efficiency, avoiding excessive truncation of critical medical information. |
Recall count (Retrieval Count) | Top 8 entries (top 8) | Ensures coverage of various potentially relevant rare disease product information, improving recall. |
Similarity threshold (Similarity Threshold) | 0.75 | Rare disease terminology is highly specific; increasing the threshold reduces interference from irrelevant content. |
Rerank result count (Rerank Count) | Top 5 entries (top 5) | Further filters retrieval results, prioritizing the most relevant product or reagent information. |
temperature | 0.1 | Consultation scenarios emphasize accuracy and factual correctness; a low temperature reduces model divergence. |
Three Common Pitfalls
- Model returns information inconsistent with reality, such as incorrect drug dosages or indications. This typically occurs due to improper chunking of relevant documents in the knowledge base, leading to truncated key information or the model failing to retrieve it effectively.
- Azure OpenAI model integration with FastGPT fails to call correctly, resulting in authentication failures or connection timeouts. This often happens because Azure OpenAI's API authentication mechanism differs from the standard OpenAI API, requiring specific configuration for
API_KEYandBASE_URL, and potentially specifyingAPI_VERSION. - Consultation results contain a large amount of irrelevant general medical knowledge, failing to focus on the rare disease product itself. This indicates that the
Similarity threshold(Similarity Threshold) is set too low, or theRecall count(Retrieval Count) is too high, causing the model to process too much generalized information.
How to Confirm Proper Configuration
- Select multiple typical rare disease product consultation scenarios. Ask questions about indications, dosage and administration, and adverse reactions. Check the accuracy and completeness of the information returned by the model.
- Use FastGPT's debugging tools to view the model's input (Prompt) and output (Completion). Confirm that the document chunks retrieved from the knowledge base are precise and cover the core content of the query.
- Query specific professional terminology or abbreviations unique to rare disease products. Verify if the model can correctly understand and provide professional, accurate explanations, checking the effectiveness of
Entity Extraction.
The values provided are common starting points. Measure them against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.