Data Characteristics
Hematologic oncology product data originates from diverse sources. These include clinical trial reports, drug monographs, medical guidelines, patent literature, and academic papers. Data updates frequently, especially clinical trial progress and new drug launch information, typically monthly or quarterly. Document structures often mix structured data and unstructured text for drug monographs and clinical trial reports. These documents contain key information such as dosage, indications, adverse reactions, mechanisms of action, and pharmacokinetics. Patent literature focuses on compound structures, preparation methods, and indication ranges. Fields are highly specific, for example, ICD-O-3 coding, disease staging, target genes, mutation types, administration routes, and pharmacodynamic indicators (e.g., CR, PR rates). Units strictly follow medical standards, such as mg/kg, μg/mL, mmol/L.
Constraints on Workflow Orchestration
The multi-source and high-frequency updates of hematologic oncology data require flexible data ingestion and integration capabilities in the workflow. The workflow must support parsing various file formats and periodic data synchronization. Mixed structured and unstructured document structures demand that the workflow accurately identifies key entities and understands contextual semantics during information extraction. This places higher demands on the effectiveness of RAG (Retrieval Augmented Generation). Highly specific fields and specialized units require precise configuration during entity recognition and knowledge graph construction to prevent information errors due to semantic ambiguity. For example, CR can mean complete remission or a specific cell type in different contexts; the workflow must distinguish this through context or predefined rules. Additionally, the ability to trace historical data is crucial. Approval and use of hematologic oncology products reference extensive prior research, so the workflow must support version management and historical snapshots to ensure traceability.
Configuration Settings
| Configuration Item | Suggested Value | Rationale |
|---|---|---|
maxContext | 6000 token | Ensures major information from clinical trial reports and drug monographs is accommodated, preventing critical context truncation. |
Recall Count | Top 8 | Given the complexity and interconnectedness of hematologic oncology knowledge, increasing the recall count covers more potentially relevant documents. |
Similarity Threshold | 0.78 | Raises the threshold to filter out documents with low relevance to hematologic oncology product inquiries, ensuring recall precision. |
Reranked Return Count | Top 3 | Focuses on the most relevant core information, reducing the burden of processing unnecessary information for the user. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Provides sufficient file parsing time when processing large clinical trial reports or multiple patent documents. |
Chunk Length | 800 characters | Balances semantic completeness and retrieval efficiency, ensuring each chunk contains enough information for contextual understanding. |
Common Mistakes
- The "Code Execution" node in the workflow fails validation when processing inputs containing "historical records." This typically occurs because the data structure of historical records does not match the code's expectations, or variable scope is misconfigured, preventing correct access or parsing of historical variables.
- The question classification node fails to accurately identify user intent for specific targeted drugs. For example, it misclassifies questions about
Imatinibas general cancer treatment. This happens because the classification model's training data does not sufficiently cover the detailed product names and mechanisms of action in hematologic oncology. - Variables defined outside a loop are inaccessible inside the loop body, leading to data processing interruptions or inconsistent results. This usually occurs when external variables are not correctly passed to the loop's scope or when variable lifecycle management is improper during workflow orchestration.
How to Verify Configuration
- Select multiple representative hematologic oncology product inquiry cases. Manually execute the workflow to verify that
RAGrecalled documents are accurate and comprehensive, and assess their relevance to user questions. - Simulate user inquiries for different types of hematologic oncology products (e.g., targeted drugs, immunotherapies, chemotherapies). Check if the question classification node correctly categorizes them and observe the accuracy of subsequent information extraction.
- Monitor workflow execution logs for
PARSE_FILE_TIMEOUT_SECONDStimeout errors ormaxContexttruncation warnings. This ensures file parsing and context handling capabilities meet expectations. - Use complex queries containing specific disease staging, gene mutations, and pharmacodynamic indicators. Verify that the workflow accurately identifies and extracts these specific fields and reflects their professional nature in the final answer.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.