Data Characteristics for This Category
Rare disease registration documents draw from diverse sources. These include clinical trial reports, pharmacological and toxicological studies, manufacturing process files, medical literature, gene sequencing reports, and approval documents from international regulatory bodies. Update frequency is relatively low, primarily occurring with new drug development milestones, clinical trial data updates, or regulatory policy changes. Document structures are complex, often containing extensive unstructured text, charts, medical terminology, gene sequence information, and biomarker data. Specific fields and units include disease classification codes (e.g., ORPHA codes), gene mutation sites, drug dosage units (e.g., mg/kg/day), biological activity units (e.g., IU/mL), and complex statistical indicators. Some documents may exist as scanned PDFs, increasing the difficulty of information extraction.
Constraints Imposed by These Characteristics on Vector Models and Indexing
The complex document structures and specialized terminology in rare disease data require vector models with advanced semantic understanding. Models must accurately capture deep textual meaning and differentiate between medical concepts. Gene sequences and biomarker data necessitate specialized embedding strategies to ensure effective indexing and retrieval of this non-textual information. Low update frequency allows for longer index rebuilding cycles. However, each update must guarantee high accuracy to avoid missing critical information. The presence of scanned PDF documents demands high precision from Optical Character Recognition (OCR) during preprocessing, directly impacting subsequent vectorization quality. Additionally, multilingual documents require vector models capable of cross-language understanding to integrate and retrieve rare disease information globally.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size | 800–1200 characters | Rare disease documents often contain long sentences and complex medical concepts; longer segments maintain contextual coherence. |
Chunk Overlap Length | 100–200 characters | Ensures semantic continuity between adjacent segments, preventing truncation of critical information. |
Recall count | 8–12 entries | Given the depth and breadth of rare disease knowledge, increasing recall items improves relevance coverage. |
Similarity threshold | 0.75–0.85 | High domain specificity requires a higher threshold to filter for highly relevant results and reduce noise. |
Max Concurrent Files | 5 | Considering the complexity of medical file parsing and potential OCR load, limiting concurrency ensures stability. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Provides sufficient parsing time for large or scanned PDF files, preventing timeout interruptions. |
Common Pitfalls
- Query results contain numerous irrelevant or generalized medical concepts. This occurs when the vector model's embedding distinction for rare disease-specific terminology is insufficient, leading to semantic generalization.
- Critical information is not retrieved during searches, resulting in missing query results. This typically stems from OCR errors in original scanned documents or a failure to effectively extract text from charts during preprocessing.
- New data is not reflected in search results after knowledge base updates. This indicates issues with index rebuilding or incremental update mechanisms, potentially leading to outdated information for users.
Verification of Configuration
- For typical rare disease cases, use core symptoms and gene mutation sites as query terms. Check if recall results include multiple relevant clinical reports and drug development progress, and evaluate their ranking relevance.
- Upload a rare disease registration PDF file containing complex tables and charts. Observe its parsing status and indexed content to ensure all critical information (e.g., dosage units, biomarkers) is correctly extracted and vectorized.
- Simulate multi-round Q&A during a rare disease new drug approval process. Evaluate the system's ability to utilize the knowledge base for complex queries and check if answers accurately cite specific passages from source documents. This helps determine if the similarity threshold is appropriate.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.