Data Characteristics in Hematologic Oncology
Hematologic oncology regulatory submission documents draw from diverse sources: clinical trial reports (Protocol, CSR), medical literature, regulatory guidelines, drug labels, and post-market safety data. Data update frequencies vary. Clinical trial data typically generates in batches after study completion, while medical literature and regulatory policies see continuous updates. Document structures are primarily unstructured text, such as lengthy clinical study reports and medical papers. These often contain numerous charts, tables, biomarker data, and genetic test results. Fields and units are highly specialized, for example, Overall Response Rate (ORR) and Progression-Free Survival (PFS). Units include days, weeks, months, years, and various biological indicator units (e.g., ng/mL, copies/mL).
Constraints Imposed by These Characteristics on Workflow Orchestration
The multimodal nature of hematologic oncology documents requires workflows to effectively process PDF or Word files containing text, tables, and image information. It also requires routing these different data modalities to appropriate processing modules. Clinical reports are extensive; a single document might exceed FastGPT's chunk_size limit. This necessitates refined document segmentation strategies to ensure information completeness and reduce recall noise. Dense specialized terminology and abbreviations demand higher precision in knowledge base recall, requiring stricter similarity thresholds. The cyclical nature of data updates means the workflow should include version management and incremental update mechanisms. This avoids reprocessing stable data and quickly integrates new clinical research advancements or regulatory requirements.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
maxContext | 8192 | Ensures capacity for complex queries and multi-turn conversations, preventing loss of critical details. |
Chunk size (Segment Length) | 800–1200 characters | Accommodates the long text reports common in hematologic oncology, balancing information completeness and recall efficiency. |
Recall count (Recall Count) | Top 8 | Given the highly specialized nature, increase recall count to cover potentially relevant information. |
Similarity threshold (Similarity Threshold) | 0.78 | Enhances recall precision, reducing interference from irrelevant or low-relevance content. |
Rerank result count (Reranked Return Count) | Top 5 | Further filters recalled results, focusing on the most relevant content and reducing model processing burden. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Accommodates the parsing time for large PDF reports, preventing processing failures due to timeouts. |
Three Common Mistakes
- When processing multimodal data and plain text simultaneously in a workflow, incorrect routing node configuration leads to plain text data being sent to a multimodal model, or vice versa, causing model input format errors.
- Setting the knowledge base search reference limit too low, such as
3000characters, prevents complete citation of critical passages in lengthy hematologic oncology reports, impacting the quality of model-generated answers. - Failure to identify table data within PDFs during document parsing results in critical clinical indicators or drug dosage information being missing from the knowledge base, leading to inaccurate or incomplete subsequent Q&A results.
How to Verify Configuration
- Upload hematologic oncology documents in various formats (PDF, Word, TXT) in a test environment. Check if knowledge base segmentation results are complete and if table and image text are correctly extracted.
- Simulate user queries for typical regulatory submission questions. Check if the workflow's returned references are accurate and complete. Compare them with original documents to verify the effectiveness of recall count and similarity thresholds.
- Monitor workflow execution logs. Confirm that
PARSE_FILE_TIMEOUT_SECONDSis sufficient to complete parsing for large files, with no timeout errors.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.