Document Parsing and Chunking for Hematologic Oncology Policies

Policy and SOP documents in hematologic oncology primarily originate from national health commission guidelines, industry association consensuses

Data Characteristics

Policy and SOP documents in hematologic oncology primarily originate from national health commission guidelines, industry association consensuses, hospital clinical pathways, drug usage specifications, and ethics review documents. These documents are updated frequently, typically annually or biennially. Some drug-related SOPs may be adjusted dynamically with new drug approvals or expanded indications. Documents often use a chapter-based structure, including introductions, definitions, diagnostic criteria, treatment plans, follow-up requirements, and complication management. Fields and units frequently involve specific medical terminology, genetic test results (e.g., mutation frequency), drug dosages (e.g., mg/kg, mg/m²), treatment courses (e.g., cycles, days), laboratory indicators (e.g., PLT, Hb), and their normal ranges.

Constraints from these Characteristics on "Document Parsing and Chunking"

The frequent updates of hematologic oncology policy documents require efficient document version management and incremental update capabilities for the knowledge base, ensuring the timeliness of recalled information. Their chapter-based structure dictates that chunking strategies must balance chapter integrity with fine-grained information extraction, preventing critical information from being split. Complex medical terms, genetic test results, and drug dosages pose higher demands on the accuracy of document parsers, requiring prevention of garbled text or semantic loss due to formatting or character set issues. Additionally, common tables, images (e.g., flow cytometry plots, chromosome karyotypes), and reference lists in documents require special handling during parsing to retain their semantic information or provide contextual links, ensuring the completeness of question answering.

Configuration Guidelines

Configuration ItemSuggested ValueRationale
Chunk size500–800 charactersBalances chapter integrity with Q&A retrieval efficiency. Avoids excessively long chunks that introduce redundancy or excessively short chunks that lose context.
Chunk Overlap Length100 charactersEnsures semantic continuity between adjacent chunks, especially in cross-paragraph process descriptions.
Parsing StrategyBy Title Splitting combined with Smart ChunkingPrioritizes respecting the document's chapter structure while using smart chunking for paragraphs without clear headings.
File Types vials Supported.pdf, .docx, .mdCovers mainstream document formats in hematologic oncology, ensuring broad compatibility.
Image tabletsOCR识别EnabledProcesses charts and images that documents may contain, such as diagnostic flowcharts or cell morphology images, ensuring image content semantics are retrievable.

Common Pitfalls

  • After uploading documents, if the large language model's output does not include image information from the document, Image tabletsOCR识别 might be disabled, or OCR results might not be effectively embedded into text chunks.
  • If clicking a reading link returned by the knowledge base results in an error message: "Only support .txt, .m", the system's default file type whitelist likely does not include common hematologic oncology document formats like .pdf or .docx. Adjust UPLOAD_FILE_MAX_SIZE related configurations.
  • After customizing chunking rules, if the large language model's answers show truncated or semantically incomplete key medical terms, Chunk size might be set too short, causing sentences or paragraphs containing complete concepts to be improperly split.

How to Verify Configuration

  • Upload a typical hematologic oncology policy document (e.g., a national treatment guideline). Check the chunk preview to ensure chapter titles and key paragraphs are fully preserved, without obvious semantic truncation.
  • For documents containing charts, images, and special medical symbols, verify that OCR results are correctly embedded into text chunks. Attempt to ask questions based on chart content to check retrieval effectiveness.
  • Select paragraphs containing complex drug dosages and genetic test results from the document. Conduct Q&A tests to confirm the model accurately extracts and answers relevant information, without garbled text or unit errors.

The values provided are common starting points. Measure them against your own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.