Data Characteristics
Quality documents in hematologic oncology originate from authoritative guidelines, clinical trial reports, drug monographs, regulatory technical review documents, and internal Standard Operating Procedures (SOPs). These documents update frequently, especially with new drug approvals, treatment regimen iterations, or regulatory policy changes. Document structures are complex, often containing extensive medical terminology, laboratory indicators, dosage units, statistical data, and charts. For example, clinical trial reports detail patient inclusion criteria, treatment regimens, adverse event grading (e.g., CTCAE v5.0), and efficacy endpoints (e.g., response rate, progression-free survival). Fields and units strictly adhere to international medical standards, such as dosage units (mg/kg, µg/dL), time units (days, weeks, months), and various biomarker detection units.
Constraints Imposed by These Characteristics on Document Parsing and Chunking
The complexity of hematologic oncology quality documents places specific demands on document parsing and chunking. High update frequency requires efficient incremental parsing and version management mechanisms to avoid reprocessing unchanged content. Nested tables and charts, especially those containing critical information like dosage, time windows, and toxicity grades, require parsers with robust structured information extraction capabilities. This ensures data context is not lost during chunking. The specialized nature of medical terminology and diverse abbreviations means that simple character- or punctuation-based chunking strategies can truncate critical information or obscure meaning. For example, a paragraph on gene mutation detection, if improperly chunked, might prevent accurate retrieval of complete mutation information. Therefore, chunking strategies must prioritize semantic integrity, focusing on the association between specialized terms and key data pairs.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
UPLOAD_FILE_MAX_SIZE | 500 MB | Hematologic oncology documents often contain high-resolution charts and extensive text, leading to large file sizes. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Complex PDF or Word documents require longer parsing times, necessitating a sufficient time window. |
Chunk size (Chunk Length) | 800–1200 characters | Balances the completeness of medical terminology with semantic coherence, preventing key information from being split. |
Chunk Overlap Length (Chunk Overlap Length) | 100–200 characters | Ensures sufficient contextual overlap between adjacent chunks, reducing information loss, especially for procedural content. |
maxContext | 32000 | High demand for model's long-text processing capability to accommodate specialized terminology and complex logic. |
CHUNK_STRATEGY | Semantic Chunking | Prioritizes semantic structures like paragraphs and headings, avoiding the splitting of medical professional concepts. |
Common Pitfalls
- Symptom: Uploading large PDF documents results in parsing failure or prolonged unresponsiveness in the backend. Reason: The
PARSE_FILE_TIMEOUT_SECONDSconfiguration is too small, not providing enough parsing time for complex documents. - Symptom: The model cannot answer questions about specific drug dosages or adverse event grades based on document content. Reason: The document parser failed to correctly identify and extract structured data from tables, leading to the loss of critical numerical information context after chunking.
- Symptom: Incomplete content parsing or corrupted formatting when processing Excel-formatted clinical data or laboratory reports. Reason: An Excel parsing plugin was not enabled or correctly configured, or the parser could not handle complex cell merging and data types.
Verification Steps
- Upload a hematologic oncology clinical trial report PDF containing complex tables and charts. Check if the parsed chunks fully retain table data and chart captions.
- Select a document with extensive medical terminology and abbreviations. Verify that the chunking results maintain the integrity of specialized terms, avoiding breaks within terms.
- Retrieve a specific drug dosage or adverse event grade from the document. Check if the recalled chunks accurately contain relevant numerical values and units, and can support the model in providing accurate answers.
The values provided are common starting points and should be measured against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.