Document Parsing and Chunking for Rare Disease Regulations

Rare disease regulation data originates from policy documents, guidelines, and expert consensuses published by national health commissions, drug

Data Characteristics

Rare disease regulation data originates from policy documents, guidelines, and expert consensuses published by national health commissions, drug administration agencies, and medical insurance bureaus. It also includes local implementation rules from various provinces and cities. These documents are typically in PDF, Word, or structured XML formats. Content covers rare disease catalogs, diagnosis and treatment pathways, medication coverage, and research ethics. Update frequency is relatively low, but once published, terms have long-term binding force. Document structures are complex, containing numerous legal provisions, medical terms, abbreviations, tables, figures, and citations. Fields and units often involve disease codes (e.g., ICD-10), generic drug names, specific treatment regimen dosage units (mg/kg), and cost codes. Extensive cross-referencing and definitions are common.

Constraints on Document Parsing and Chunking

Complex and irregular document structures, especially nested clauses and numerous tables, demand high accuracy in document parsing. The parser must identify different levels of headings, paragraphs, lists, and tables, correctly extracting their content to prevent information loss or misalignment. Dense medical terminology and abbreviations can lead to inaccurate tokenization, impacting subsequent semantic understanding and retrieval. The low document update frequency means initial parsing is labor-intensive, but later maintenance costs are relatively controllable. The focus is on ensuring parsing completeness and standardization. Cross-references between clauses require considering contextual relevance during chunking to avoid fragmenting critical information. Specific fields and units, such as dosage units, require the parser to distinguish numbers from units to ensure value correctness.

Configuration Settings

Configuration ItemRecommended ValueRationale
UPLOAD_FILE_MAX_SIZE500 MBRare disease policy documents may contain many scanned images or high-resolution images, leading to large file sizes.
PARSE_FILE_TIMEOUT_SECONDS600 secondsParsing large PDF or Word documents can be time-consuming; this prevents timeout interruptions.
Chunk size (Chunk Length)800–1200 charactersParagraphs in rare disease regulation documents are typically long, containing complete policy clauses or medical descriptions.
Chunk overlap (Chunk Overlap)100–200 charactersEnsures contextual continuity at chunk boundaries, preventing important information from being truncated.
Table Parsing StrategyStructured ExtractionRegulation documents contain many tables; accurate extraction of table content and semantic relationships is necessary.
Embedding Modeltext-embedding-ada-002Suitable for processing specialized domain texts, with good understanding of medical terminology and policy provisions.

Common Pitfalls

  • Uploading large PDF documents results in a timeout of 360000ms exceeded error: This usually indicates that PARSE_FILE_TIMEOUT_SECONDS is set too low and does not cover the actual time required for document parsing.
  • After uploading some Word documents, table content is not correctly chunked or is missing: The complex table structures embedded in the document were not recognized by the parser, leading to incomplete content extraction.
  • Question-answering results do not fully quote the "answer" from the original knowledge base text: The knowledge base chunking granularity is too fine, causing the complete context of the original text to be split. The model fails to retrieve the full "answer" context.

Verification Steps

  • Select representative rare disease regulation documents. Upload them and check if all text content, especially long paragraphs and table information, is fully preserved in the knowledge base.
  • Conduct question-answering tests on key clauses, definitions, or specific medical terms within the documents. Verify that retrieval results are accurate and include the original text.
  • Examine the length and overlap of chunks in the knowledge base. Ensure each chunk contains sufficient context and that critical information is not fragmented.
  • Use FastGPT's knowledge base preview feature. Randomly select multiple chunks and manually verify that their content is semantically complete and untruncated.

The values provided are common starting points. Measure them against your own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.