Data Characteristics
Data for hematologic oncology products and reagents primarily originates from clinical research reports, drug inserts, diagnostic reagent kit instructions, NMPA (National Medical Products Administration) approval documents, and related regulatory files. These documents undergo frequent updates, especially with new drug approvals, expanded indications, or revised instructions. Document structures are complex, often containing extensive medical terminology, charts, experimental data (e.g., flow cytometry results, gene sequencing reports), clinical trial outcomes, adverse event lists, and dosage and administration information. Fields and units are highly specialized, including dosage units (mg/kg, IU), concentration units (µg/mL, nM), time units (weeks, cycles), and specific tumor response evaluation criteria (e.g., CR, PR).
Constraints on Document Parsing and Chunking
The complexity of hematologic oncology documents places high demands on document parsing. Extensive specialized terminology and abbreviations require parsers to accurately identify and understand context, preventing incorrect tokenization or loss of critical information. Charts and experimental data often appear as images, necessitating OCR technology for text extraction and correct contextual association. Frequent document updates mean the system must support rapid incremental parsing and version management. Furthermore, critical numerical information such as dosage and concentration is typically accompanied by units; parsing must bind values with their units to ensure information completeness and prevent unit loss or incorrect matching that could lead to comprehension errors. These constraints dictate that chunking strategies must balance semantic integrity and information granularity.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size | 500–800 characters | Balances paragraph semantic integrity with retrieval efficiency, preventing overly long chunks from diluting key information or overly short chunks from losing context. |
Chunk Overlap Length | 80–120 characters | Ensures contextual continuity between adjacent chunks, particularly for content spanning pages or sections. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Provides sufficient parsing time for large PDFs or Excel files with many images, preventing parsing timeouts. |
maxContext | 3000 Tokens | Accommodates the terminology-dense and context-dependent nature of hematologic oncology, ensuring the model receives enough information. |
IMAGE_OCR_ENABLED | True | Clinical reports and reagent instructions often contain charts; enabling OCR extracts key text information from images. |
TABLE_PARSING_STRATEGY | markdown_table | Ensures experimental data and adverse event tables are parsed structurally, facilitating subsequent retrieval. |
Common Pitfalls
PARSE_FILE_TIMEOUT_SECONDSerrors occur when parsing large PDF documents because insufficient parsing time is allocated for complex documents.- Relevant information is not retrieved after uploading flow cytometry images or gene sequencing reports in image format because
IMAGE_OCR_ENABLEDis not set toTrue, preventing text recognition within images. - Values and units are separated or formatting is chaotic after parsing dosage tables in Excel because an unsuitable
TABLE_PARSING_STRATEGYis chosen for structured parsing.
Verification Steps
- Select several typical hematologic oncology product inserts (PDF format, including charts and tables), upload them, and check that parsing status is normal without timeouts or failures.
- Perform keyword searches on the parsed document content to verify that key information, including text from images or table data, can be accurately retrieved.
- Examine the parsed chunks to ensure each chunk has semantic integrity and that no critical information is truncated or context is lost.
The values provided are common starting points and should be measured against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.