Data Characteristics
Rare disease product and reagent data originates from clinical trial reports, drug monographs, academic papers, international rare disease databases (e.g., Orphanet, OMIM), and regulatory documents. Data updates are infrequent, typically occurring periodically with new drug approvals, clinical research advancements, or guideline revisions. However, updates are more frequent for cutting-edge fields like gene therapy. Document structures commonly include disease definitions, epidemiology, genetic background, diagnostic criteria, treatment plans, drug dosages, side effects, and interactions. Fields often contain extensive medical terminology, gene sequences, protein names, and complex chemical formulas. Units involve measurement units (e.g., mg/kg), time units (e.g., weeks, months), and biological units (e.g., IU/mL), frequently accompanied by upper and lower range descriptions.
Constraints on Knowledge Base Retrieval and Recall
The infrequent update rate of rare disease data means significant effort is required for initial data collection and cleaning. Subsequent maintenance costs are manageable, focusing on tracking key literature releases. Complex document structures and dense specialized terminology demand advanced text segmentation and embedding model selection. This ensures semantic integrity and prevents critical information loss due to truncation. Unstructured or semi-structured information, such as gene sequences and chemical formulas, requires the knowledge base to handle special characters and long sequences. Pure text retrieval may not capture deep associations. Precision is critical for fields like drug dosages and side effects. Retrieval results must be highly accurate; any deviation can impact consultation quality. Understanding multiple units and range descriptions requires the retrieval system to recognize and match the same concept expressed in different forms.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size | 800–1200 characters | Rare disease documents have long paragraphs containing multi-dimensional information; this ensures semantic integrity within a single segment. |
Chunk Overlap Length | 100–200 characters | Reduces information loss at segment boundaries, improving contextual continuity. |
Recall count | 8–12 entries | Rare disease information is dense; increasing recall items improves coverage. |
Similarity threshold | 0.75–0.85 | Ensures relevance of retrieval results, reducing interference from irrelevant information. |
Rerank result count | 3–5 entries | Selects the most relevant few items for the large model, preventing information overload. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Handles large PDFs or documents with complex charts, preventing parsing timeouts. |
Common Pitfalls
- Knowledge base image output addresses are truncated. This usually results from file storage paths or URL lengths exceeding system limits, leading to incomplete links.
- Imported knowledge base content is bilingual (Chinese and English), but the large model fails to read corresponding content during responses. This may occur if the embedding model's semantic understanding of mixed Chinese and English text is insufficient, or if the segmentation strategy splits bilingual content.
- During testing of a question-classification-based workflow, only the first question has a knowledge base reference, while subsequent questions do not. This could stem from limitations in context passing within the workflow configuration or subsequent questions failing to trigger knowledge base retrieval conditions correctly.
Verification Steps
- Upload documents containing complex medical terminology, gene sequences, and dosage information. Check if knowledge base segmentation is reasonable and ensures critical information is not truncated.
- Test the relevance of knowledge base recall results for specific rare disease product queries. Verify accurate matching of key fields such as drug dosages and side effects.
- Simulate user consultation scenarios. Ask questions with mixed Chinese and English or specialized terminology. Validate that the large model correctly references corresponding content from the knowledge base.
- Check system logs to confirm no exceptions occurred during knowledge base retrieval, such as file parsing timeouts (
PARSE_FILE_TIMEOUT_SECONDS) or storage path errors.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.