Data Characteristics
Infectious disease protocols and SOP documents originate from medical institutions, disease control centers, and health administrative departments. These documents update frequently due to national policy adjustments, treatment guideline updates, or the emergence of new infectious diseases. Documents are typically hierarchical PDFs or Word files. They contain medical terminology, pathogen names, diagnostic criteria, treatment plans, and isolation measures. Common fields include disease codes (e.g., ICD-10), drug dosage units (mg/kg), time periods (days, hours), and specific biosafety levels (BSL-2, BSL-3). Documents often include flowcharts, tables, and appendices detailing operating procedures or providing references.
Constraints on Document Parsing and Chunking
The hierarchical structure and high density of specialized terminology in infectious disease protocol documents require precise identification of sections, headings, and body text to avoid semantic fragmentation. Frequent updates necessitate efficient incremental parsing mechanisms to ensure knowledge base timeliness. Flowcharts and tables challenge traditional text parsing, potentially leading to information loss or structural disorganization. Precise information, such as disease codes and dosage units, must maintain integrity during chunking to prevent loss of contextual relevance. These constraints mean chunking must consider text length, semantic completeness, and structural preservation to meet the accuracy requirements for subsequent RAG retrieval.
Configuration Settings
| Configuration Item | Suggested Value | Rationale |
|---|---|---|
chunk_overlap | 50–100 characters | Ensures contextual continuity at chunk boundaries, especially for specialized terminology and process descriptions. |
chunk_size | 800–1200 characters | Balances retrieval granularity with contextual completeness, accommodating the detailed descriptions in infectious disease documents. |
max_page_tokens | Calibrate by measurement | Prevents single-page processing timeouts or memory overflows for complex PDFs with charts or scanned images. |
embedding_model | text-embedding-ada-002 | Balances accuracy and cost, offering good understanding of medical terminology. |
parser_config.split_by_title | true | Uses the document's hierarchical title structure for chunking, preserving chapter semantic integrity. |
parser_config.max_table_rows | 20 | Limits the number of table rows parsed to prevent oversized chunks or parsing failures from very large tables. |
Common Pitfalls
- Uploading a PDF file results in empty content. This may occur if the PDF is a scanned image or an encrypted document, causing text extraction to fail.
- Non-structured code text, such as Java API documentation, is imported but yields empty parsing results. The parser defaults to natural language processing and cannot recognize code structures.
- After uploading documents via API, a prolonged "Indexing" status often indicates that the document is too large or contains complex elements, leading to extended backend parsing times or task stagnation.
Verification Steps
- Randomly select parsed documents. Preview chunk content through the knowledge base management interface. Check if each chunk is semantically complete and without obvious truncation.
- For documents containing flowcharts and tables, verify that key information is correctly extracted and included in the corresponding chunks.
- For frequently updated documents, upload new versions. Check if incremental parsing completes quickly and updates the knowledge base content without affecting the availability of existing chunks.
The values provided are common starting points. Measure them against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.