Data Characteristics
Rare disease quality documentation data is highly dispersed. Sources include rare disease drug review guidelines from various national regulatory agencies, orphan drug designation criteria, clinical trial protocols, pharmacovigilance reports, and clinical evidence from medical journals on rare disease drug treatments. Data update frequency is relatively low, typically occurring with new drug approvals or guideline revisions, which can take months or even years. Document structures are primarily unstructured text, such as PDF regulatory documents and Word clinical reports. These documents contain extensive specialized medical terminology, gene sequence information, drug molecular formulas, and dosage units. Field and unit specificities include precise descriptions of rare disease diagnostic criteria (e.g., disease phenotype scores), drug mechanisms of action (e.g., enzyme activity units U/mL), and patient population characteristics (e.g., specific gene mutation types).
Constraints on Deployment and Upgrade
The dispersed nature of rare disease quality documentation data requires multi-source heterogeneous data integration capabilities during deployment. This necessitates support for uploading and parsing various file formats. The low update frequency means a large initial data volume for knowledge base construction, but less pressure for subsequent incremental updates. Deployment solutions should prioritize efficient import of large initial datasets. The unstructured nature and high density of specialized terminology in documents demand more sophisticated text segmentation strategies and embedding models. Generic models may not effectively capture rare disease-specific semantic associations. The specificity of fields and units, especially involving gene sequences and molecular formulas, may invalidate default text cleaning and entity recognition rules. This requires custom preprocessing pipelines during deployment. Furthermore, rare disease data often involves patient privacy, so the deployment environment must meet strict data security and compliance requirements.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
UPLOAD_FILE_MAX_SIZE | 500 MB | Rare disease regulatory documents and clinical trial reports can be large; ensure full document upload. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Longer parsing times are needed for large PDFs or documents with complex tables to avoid timeouts. |
Chunk size (Segment Length) | 800–1200 characters (characters) | Rare disease documents are highly specialized and context-dependent; longer segments maintain semantic integrity. |
Similarity threshold (Similarity Threshold) | 0.78 | Increase the threshold to ensure high relevance of retrieved results to rare disease queries, reducing inaccurate medical information. |
maxContext | 8192 | A longer context window helps the model understand complex rare disease pathology and drug mechanisms. |
Rerank result count (Reranked Results) | Top 10 entries (top 10) | After fine-grained ranking, provide enough relevant snippets for user reference to improve information coverage. |
Common Pitfalls
- Crashes on chat or knowledge base pages: This indicates insufficient deployment environment resources or model loading failures. For example, running large models like
deepseek-32bon hardware with inadequate configuration can lead to memory overflow or exhaustion of computing resources. - Raw tags like
<think></think>appear in model responses: This means the model's output post-processing logic is not correctly configured. This usually happens when interceptors or output parsing scripts are not properly set up during deployment, causing the model's raw output to be presented directly to the user without processing. - Long unresponsiveness or errors when uploading large documents: This is due to
UPLOAD_FILE_MAX_SIZEorPARSE_FILE_TIMEOUT_SECONDSbeing set too low. Rare disease documents often contain numerous charts and complex structures, and default configurations are insufficient for their parsing requirements.
Verification Steps
- Upload a PDF rare disease drug instruction manual containing gene sequences or complex molecular formulas. Verify that the file is successfully parsed and segmented.
- Query about diagnostic criteria or treatment plans for a specific rare disease (e.g., Duchenne muscular dystrophy). Check if the model's answer accurately cites relevant document snippets from the knowledge base and validate the effect of the
Similarity threshold. - Simulate multiple users accessing and querying simultaneously. Observe system response speed and resource utilization to ensure system stability under expected concurrency. This verifies the performance of parameters like
maxContextunder actual load.
Note: The values provided are common starting points. They should be measured against specific samples and requirements.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.