Data Characteristics in Hematologic Oncology
Data for hematologic oncology clinical trial pre-screening comes from diverse sources. These include Electronic Health Records (EHR), genetic testing reports, imaging results, pathology reports, and treatment history. EHR data updates frequently, potentially daily or in real-time. Genetic testing and pathology reports are typically one-time generations but may be reissued during treatment.
EHRs are mostly semi-structured text, containing fields like diagnosis, medication, and lab results. Genetic reports include structured information such as specific gene mutations, fusion genes, and copy number variations, often with interpretive text. Specific fields include leukocyte count, hemoglobin, and platelet levels, along with bone marrow biopsy results, cytogenetic karyotypes, and FISH test results. Units often include g/L, ×10^9/L, and %.
Constraints Imposed by Data Characteristics on Workflow Orchestration
The multi-source and heterogeneous nature of hematologic oncology data requires robust data integration and cleaning capabilities in the workflow. The semi-structured nature of EHRs necessitates complex entity recognition and information extraction during data ingestion. For example, accurately extracting tumor staging and treatment plans from free text. Structured information in genetic reports requires precise parsing to match pre-screening criteria for gene mutations or expression status.
Differences in data update frequency mean the workflow must support incremental updates and periodic full synchronizations. This ensures the timeliness and accuracy of pre-screening results. The presence of specific fields, such as blast cell percentage in bone marrow, requires the workflow's semantic understanding model to recognize and correctly process these domain-specific terms. Additionally, various data report formats (e.g., PDF, images, text) require the workflow to support multi-format processing during file parsing. It must also unify data from different sources into a standardized data model to prevent data loss or misinterpretation.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Segment Length) | 500–800 characters | Hematologic oncology medical records and reports often contain long descriptive paragraphs. This length helps preserve contextual integrity and prevents truncation of critical information. |
Recall count (Recall Count) | Top 10–15 items | Patient information and genetic reports have complex associations. Increasing the recall count improves coverage and ensures no potential matching information is missed. |
Similarity threshold (Similarity Threshold) | 0.75–0.85 | Clinical trial conditions typically require precise matching for key indicators. A high threshold reduces false positives and lowers the risk of ineligible patients entering subsequent processes. |
maxContext | 8000–12000 tokens | When integrating multiple patient reports (EHR, genetic, pathology), a larger context window is needed to hold all relevant information for comprehensive evaluation. |
PARSER_CONCURRENCY_LIMIT | 2–4 | Report parsing in hematologic oncology involves OCR and text extraction, which are computationally intensive. Limiting concurrency prevents system overload. |
TOOL_RESPONSE_TIMEOUT_SECONDS | 60–90 seconds | External genetic database queries or clinical guideline lookups can be time-consuming. Extending the tool response timeout prevents task interruptions. |
Common Pitfalls
- Workflow execution timeout with
TOOL_CALL_TIMEOUT: This occurs when external gene database queries or clinical guideline APIs respond slowly, but the tool call timeout is set too short. - Omission or error in key indicators (e.g., gene mutation sites) in pre-screening results: This happens when the entity recognition model for unstructured text during data ingestion is insufficiently trained, leading to inaccurate extraction of critical medical terms.
- Workflow fails to process new pathology report formats, causing data parsing errors: This indicates that the file parser is not updated or not configured to handle OCR for new PDF or image report types.
Validation Steps
- Select a batch of patient data known to meet and not meet pre-screening criteria. Run the workflow and verify if the output matching results align with expectations. Calculate accuracy and recall rates.
- Check workflow logs to confirm all external tool calls (e.g., gene database queries, clinical guideline lookups) execute successfully. Ensure no
ERRORorWARNINGlevel timeouts or connection failures are recorded. - Randomly sample multiple raw data points from different sources (EHR, genetic reports, pathology reports). Compare them with the structured intermediate data processed by the workflow. Confirm that key fields, units, and textual descriptions are accurately extracted and standardized.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.