Data Characteristics
Rare disease regulation data comes from official sources. These include regulations, guidelines, directories, and notices from the National Medical Products Administration (NMPA), the National Health Commission, and local medical insurance bureaus. Documents are typically PDFs, with some in HTML. Update frequencies vary. National laws and regulations may have major revisions every few years. Local medical insurance directories or clinical guidelines may update annually.
Document structures are chapter- and item-based. They contain many specialized terms, disease classification codes (e.g., ICD-10), generic drug names, indications, and reimbursement scopes. The data also includes medical terminology, dosage units (mg/kg), and administration frequencies. These are critical for understanding the regulatory text.
Constraints on Citation and Traceability
Rare disease regulatory documents are official and legally binding. This makes citation accuracy and traceability critical. Nested tables and images in PDF documents may lose context during chunking, affecting recall quality.
Uncertain update frequencies require knowledge bases to have version management. This ensures citations refer to the latest effective regulations. Recognizing specialized terms like disease classification codes and generic drug names requires optimizing text embedding models or configuring specialized dictionaries to improve recall precision.
Dosage units and administration frequencies require precise presentation in Q&A to avoid misinformation. This requires citation snippets to accurately capture original details and support highlighting.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Chunk Size) | 400–600 characters | Rare disease regulatory texts are logically rigorous. Chunks that are too long introduce irrelevant information. Chunks that are too short lack sufficient context. |
Recall count (Recall Count) | Top 5 | Regulatory Q&A demands high accuracy. Increasing recall count helps cover more relevant regulations while avoiding redundancy. |
Similarity threshold (Similarity Threshold) | 0.75 | This ensures recalled snippets are highly relevant to the user's query, reducing the risk of incorrect citations. |
Rerank result count (Rerank Return Count) | Top 3 | This further refines recalled results. It prioritizes the most matching regulatory clauses, improving citation quality. |
ENABLE_PDF_OCR | true | Many regulatory documents are published as scanned PDFs. Enabling OCR ensures text can be extracted. |
EMBEDDING_MODEL_NAME | text-embedding-ada-002 | This model is suitable for processing medical terminology and regulatory texts, improving semantic matching accuracy. |
Common Misconfigurations
- No citation sources return after a user query. The
Similarity threshold(Similarity Threshold) may be set too high, filtering out valid passages with slightly lower relevance. - Returned citation sources have low relevance to the user's question. The
Chunk size(Chunk Size) may be too long. Individual chunks contain too much irrelevant information, diluting key semantics. - During knowledge base testing, some table data in regulatory texts is not cited. This occurs if
ENABLE_PDF_OCRis not enabled or the OCR engine has insufficient table recognition capabilities, preventing table content from being indexed.
Verification Steps
- Query reimbursement policies or medication regulations for specific diseases listed in the "Rare Disease Directory" published by the National Medical Products Administration. Check if the returned citation sources point to the latest version of the document and specific clauses.
- Select 10-15 regulatory clauses containing specialized medical terms and dosage units. Construct queries to test them. Verify that citation snippets completely and accurately capture these key details. Adjust
Similarity threshold(Similarity Threshold) to ensure recall. - Upload a rare disease clinical guideline PDF containing complex tables and images. Ask questions related to the table data. Verify that text content within tables can be accurately cited. Check the actual effect of
ENABLE_PDF_OCR.
The values provided are common starting points. Measure them against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.