Document Parsing and Chunking for Infectious Disease Quality Documents

Quality documents in the infectious disease domain draw from diverse sources. These include diagnostic guidelines, prevention and control plans, and

Data Characteristics

Quality documents in the infectious disease domain draw from diverse sources. These include diagnostic guidelines, prevention and control plans, and technical operation specifications issued by national health commissions and various levels of disease control centers. They also include hospital-internal infection control manuals, emergency plans, and case report templates. These documents update frequently, especially with new or emergent infectious diseases, where guidelines may iterate within weeks. Document structures are complex, containing pure text descriptions, numerous tables (e.g., pathogen detection results, antimicrobial susceptibility test data), flowcharts (e.g., diagnostic processes, isolation measures), and images (e.g., microscopic images of microorganisms, pathological sections). Common fields and units include Pathogen Name, ICD-10 Code, Antibiotic Name, Dosage Unit (mg, g), Time Unit (hours, days), and Microbial Count (CFU/mL).

Constraints on Document Parsing and Chunking

The high update frequency of infectious disease quality documents requires rapid processing and incremental update capabilities for the document parsing system to ensure knowledge base timeliness. Complex document structures, particularly tables and flowcharts, challenge traditional text chunking methods. Chunking directly by character or paragraph can truncate table content, leading to a loss of contextual relevance. Key step descriptions within flowcharts may also be misinterpreted or overlooked. Additionally, specific fields and units, such as ICD-10 Code or CFU/mL, must be accurately identified and retained for subsequent retrieval and question answering. Inaccurate parsing can compromise the integrity of these specialized terms, directly impacting the accuracy and professionalism of question answering. Therefore, the document parsing and chunking process needs to balance content completeness, structural accuracy, and specialized terminology recognition.

Configuration Guidelines

Configuration ItemRecommended ValueRationale
Chunk Length800–1200 charactersBalances contextual completeness with the information density of a single chunk, preventing excessive splitting of a single knowledge point.
Chunk Overlap Length100–200 charactersEnsures contextual continuity between chunks, especially when processing process descriptions or detailed guidelines.
UPLOAD_FILE_MAX_SIZE500 MBAccommodates large guideline documents containing numerous charts or scanned images.
PARSE_FILE_TIMEOUT_SECONDS600 secondsHandles large PDF or DOCX files, preventing parsing timeouts.
Retain Table StructureEnableTable data in infectious disease documents carries crucial information; its integrity must be maintained.
Recall CountTop 5Ensures comprehensive recall results for question answering, covering multiple angles of information.

Common Pitfalls

  • After document parsing, some table data in the knowledge base cannot be correctly recalled during questioning, manifesting as missing key numerical values in the answer. This occurs when table content is incorrectly truncated by row or character, leading to incomplete semantics and affecting retrieval matching.
  • Uploading a large PDF document results in a long period of unresponsiveness during parsing, eventually displaying a Parsing Timeout error. This happens when PARSE_FILE_TIMEOUT_SECONDS is not set sufficiently for large files, or the file size exceeds the system's processing capacity.
  • During knowledge base question answering, queries about specific disease diagnostic processes lack logical connections between steps in the recall results. This may be due to insufficient consideration of the integrity of flowcharts or step descriptions during document chunking, leading to critical process nodes being dispersed across different chunks.

Verification Steps

  • Select several representative infectious disease diagnostic guidelines or prevention and control plans. Upload and parse them. Randomly query table data or process steps from these documents and check if the answers accurately and completely present the original information.
  • Inspect parsed documents in the knowledge base by manually browsing their chunked content. Focus on whether tables, figure captions, and key specialized terms are fully retained within a single chunk without unnatural truncation.
  • Upload and parse documents of varying sizes (e.g., 5MB, 50MB, 100MB) and monitor parsing time. Ensure that most documents complete parsing successfully within the set PARSE_FILE_TIMEOUT_SECONDS, with an acceptable error rate.

The values provided are common starting points. Measure against your own samples to determine the most suitable configuration.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.