Data Characteristics in this Category
Rare disease data originates from clinical research reports, drug inserts, genetic testing reports, academic papers, and patient registry systems. These documents update infrequently, primarily annually, coinciding with new drug approvals, clinical trial results, or guideline revisions. Document structures are complex, often containing extensive medical terminology, abbreviations, and specific table formats. Fields and units are highly specialized, for example, dosage units (mg/kg/day), gene loci (exon numbers, mutation types), and clinical phenotype descriptions (ICD-10 codes, HPO terms). Data frequently intersperses unstructured text, such as disease progression descriptions and treatment plan details, with structured data like genetic test result tables.
Constraints from these Characteristics on "Document Parsing and Chunking"
The complexity of rare disease documents poses parsing challenges. Embedded tables, chart descriptions, and scanned content require high-accuracy OCR recognition and table structure restoration. The density of specialized terminology and abbreviations demands that the tokenizer correctly identifies medical entities, avoiding incorrect splitting or omissions. Low update frequency means knowledge base construction must emphasize historical version management to ensure information traceability. The specialized nature of fields and units requires retaining contextual semantics after chunking. For example, drug dosage must associate with administration routes and target populations; simple fixed-length truncation is insufficient. When long documents contain information on multiple diseases, genes, or drugs, effectively distinguishing and creating independent, semantically complete knowledge chunks for each entity is critical for improving recall accuracy.
Configuration Settings
| Configuration Item | Recommended Value | Rationale for this Value |
|---|---|---|
chunk_size | 800–1200 characters | Ensures individual knowledge chunks contain sufficient contextual semantics, accommodating the long-text descriptive nature of the rare disease domain. |
overlap_size | 100–200 characters | Minimizes information loss at knowledge chunk boundaries, especially when describing disease mechanisms or treatment plans. |
ocr_enabled | true | Processes non-text information common in the rare disease domain, such as scanned literature and image reports. |
table_parsing_mode | strict | Precisely extracts and restores table data structures from clinical research reports and genetic testing reports. |
max_file_size_mb | 500 MB | Accommodates large file uploads, such as extensive clinical trial reports or merged academic papers. |
timeout_seconds | 600 seconds | Handles potentially long parsing times for complex documents (e.g., PDFs with numerous tables and images). |
Three Common Mistakes
- Uploading large PDF files results in "parsing failed" or "file too large" messages. This occurs because
max_file_size_mbortimeout_secondsparameters are set too low, causing the file to exceed memory or time limits during parsing. - Knowledge base query results lack or have incomplete information about specific gene mutations. This often happens when
ocr_enabledis not set totrueduring document parsing, preventing critical information in scanned genetic test reports from being recognized. - When Excel-format clinical data tables are uploaded, vectorization indexing is incorrect, leading to poor recall for relevant queries. This typically indicates that
table_parsing_modeis not set tostrict, failing to correctly identify and extract header-to-cell associations.
How to Verify Proper Configuration
- Upload a rare disease clinical research report PDF containing complex tables and scanned pages. Check if parsed knowledge chunks fully retain table structures and text content from images.
- For a research paper with various rare disease terms and abbreviations, use keyword queries to verify that knowledge chunks for different terms are accurately recalled and maintain complete contextual semantics.
- Upload an Excel file of a large rare disease dataset exceeding
200 MB. Confirm that the parsing process completes without errors and that structured data within it can be effectively queried.
The values provided are common starting points and should be measured against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.