Data Characteristics in This Category
Rare disease quality documents come from diverse sources. These include drug inserts, clinical trial reports, regulatory submission materials (e.g., CTD modules), post-market safety reports, and orphan drug designation documents from various countries. Update frequency is relatively low, primarily occurring with new drug approvals, expanded indications, or significant safety information updates. Document structures often contain extensive medical terminology, abbreviations, dosage units (e.g., mg/kg, IU), and specific disease classification codes (e.g., ORPHAcode, ICD-10). Files are predominantly PDFs, frequently including tables, charts, and scanned images. Field names are highly specialized, for example, PK parameters, pharmacodynamic indicators, adverse event classification.
Constraints on Knowledge Base Retrieval and Recall
The dense specialized terminology and abbreviations in rare disease documents demand specific text understanding models and tokenization strategies. These strategies must correctly identify and index proper nouns. The presence of dosage units and disease codes makes precise matching critical; fuzzy matching can lead to misinterpretations. The PDF format, including tables and scanned images, challenges file parsing capabilities, potentially resulting in incomplete text extraction or formatting errors. Low update frequency means historical data is highly important. The knowledge base needs to support efficient management and retrieval of historical versions. Information is often dispersed across different document types. Building a knowledge graph or using metadata for retrieval enhancement improves recall accuracy.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk Size | 500–800 characters | Rare disease documents have high content density. Information within paragraphs is highly correlated. This avoids excessive chunking that leads to semantic loss. |
Chunk Overlap | 100–150 characters | Ensures contextual continuity, especially when specialized terms and data descriptions span chunk boundaries. |
Text Understanding Model | Select an embedding model that supports medical domain vocabulary | Improves understanding of rare disease-specific terminology and abbreviations, reducing semantic drift. |
Recall Count | Top 5–8 items | Rare disease information is highly specialized. Recalling more relevant snippets initially improves the accuracy of subsequent re-ranking. |
Similarity Threshold | 0.75–0.85 | Ensures high relevance between recall results and query intent, filtering out irrelevant information. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Accommodates the time required to parse large PDF files or documents containing complex tables and charts. |
Three Common Mistakes
- Knowledge base query returns incomplete results or misses critical information: This occurs when the original PDF file parsing fails. Some content (e.g., text from scanned images or table data) is not extracted correctly, leading to an incomplete index.
- The
Text Understanding Modeldropdown is empty when creating a knowledge base: This happens when the indexing model is not configured correctly or fails to connect to the model service. The system cannot load available embedding models. - Queries for specific disease codes or drug dosages recall a large amount of irrelevant information: This is due to the tokenizer failing to correctly identify medical proper nouns and units. It splits them into common words, resulting in overly granular indexing or semantic ambiguity.
How to Confirm Correct Configuration
- Upload a typical rare disease drug insert (e.g., a PDF containing
ORPHAcodeandmg/kgdosages). Check file parsing logs to confirm noPDF Parse ErrororTimeouterrors. Verify the completeness of the extracted text content. - Use the knowledge base's search test function. Input queries containing rare disease names, specific gene mutations, or key clinical indicators. Observe if the recall results include expected document snippets and evaluate their relevance.
- For uploaded documents, use the
Knowledge Base Overviewto check if the number of chunks and average chunk length match the expected configuration. This ensures the chunking strategy is effectively implemented. - Simulate user questions, such as querying the latest treatment plan or adverse reactions for a rare disease. Verify that the knowledge points cited in the AI's answer accurately originate from the knowledge base and contain no obvious factual errors.
The values provided are common starting points. Measure them against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.