Knowledge Base Retrieval and Recall for Rare Disease Products

Rare disease product and reagent data originates from clinical trial reports, drug monographs, academic papers, international rare disease databases

Data Characteristics

Rare disease product and reagent data originates from clinical trial reports, drug monographs, academic papers, international rare disease databases (e.g., Orphanet, OMIM), and regulatory documents. Data updates are infrequent, typically occurring periodically with new drug approvals, clinical research advancements, or guideline revisions. However, updates are more frequent for cutting-edge fields like gene therapy. Document structures commonly include disease definitions, epidemiology, genetic background, diagnostic criteria, treatment plans, drug dosages, side effects, and interactions. Fields often contain extensive medical terminology, gene sequences, protein names, and complex chemical formulas. Units involve measurement units (e.g., mg/kg), time units (e.g., weeks, months), and biological units (e.g., IU/mL), frequently accompanied by upper and lower range descriptions.

Constraints on Knowledge Base Retrieval and Recall

The infrequent update rate of rare disease data means significant effort is required for initial data collection and cleaning. Subsequent maintenance costs are manageable, focusing on tracking key literature releases. Complex document structures and dense specialized terminology demand advanced text segmentation and embedding model selection. This ensures semantic integrity and prevents critical information loss due to truncation. Unstructured or semi-structured information, such as gene sequences and chemical formulas, requires the knowledge base to handle special characters and long sequences. Pure text retrieval may not capture deep associations. Precision is critical for fields like drug dosages and side effects. Retrieval results must be highly accurate; any deviation can impact consultation quality. Understanding multiple units and range descriptions requires the retrieval system to recognize and match the same concept expressed in different forms.

Configuration Settings

Configuration ItemRecommended ValueRationale
Chunk size800–1200 charactersRare disease documents have long paragraphs containing multi-dimensional information; this ensures semantic integrity within a single segment.
Chunk Overlap Length100–200 charactersReduces information loss at segment boundaries, improving contextual continuity.
Recall count8–12 entriesRare disease information is dense; increasing recall items improves coverage.
Similarity threshold0.75–0.85Ensures relevance of retrieval results, reducing interference from irrelevant information.
Rerank result count3–5 entriesSelects the most relevant few items for the large model, preventing information overload.
PARSE_FILE_TIMEOUT_SECONDS600 secondsHandles large PDFs or documents with complex charts, preventing parsing timeouts.

Common Pitfalls

  • Knowledge base image output addresses are truncated. This usually results from file storage paths or URL lengths exceeding system limits, leading to incomplete links.
  • Imported knowledge base content is bilingual (Chinese and English), but the large model fails to read corresponding content during responses. This may occur if the embedding model's semantic understanding of mixed Chinese and English text is insufficient, or if the segmentation strategy splits bilingual content.
  • During testing of a question-classification-based workflow, only the first question has a knowledge base reference, while subsequent questions do not. This could stem from limitations in context passing within the workflow configuration or subsequent questions failing to trigger knowledge base retrieval conditions correctly.

Verification Steps

  • Upload documents containing complex medical terminology, gene sequences, and dosage information. Check if knowledge base segmentation is reasonable and ensures critical information is not truncated.
  • Test the relevance of knowledge base recall results for specific rare disease product queries. Verify accurate matching of key fields such as drug dosages and side effects.
  • Simulate user consultation scenarios. Ask questions with mixed Chinese and English or specialized terminology. Validate that the large model correctly references corresponding content from the knowledge base.
  • Check system logs to confirm no exceptions occurred during knowledge base retrieval, such as file parsing timeouts (PARSE_FILE_TIMEOUT_SECONDS) or storage path errors.

The values provided are common starting points and should be measured against the reader's own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.