Data Characteristics for This Category
Solid tumor quality documentation comes from various sources. These include clinical trial reports, pathology diagnostic reports, surgical records, genetic testing reports, and Good Manufacturing Practice (GMP) documents. Update frequencies vary: clinical trial data might update with phased reports, while GMP files might revise annually or due to regulatory changes. Document structures are mostly unstructured or semi-structured text. They contain extensive medical terminology, abbreviations, and specific table formats. For example, pathology reports include fields like tumor type, grade, and margin status. Genetic testing reports list gene mutation sites and variant frequencies. Field values often include specific units such as millimeters (mm), micromoles (µmol/L), and copy numbers (copies). Document languages are primarily Chinese and English, mixing professional medical vocabulary and standardized expressions.
Constraints Imposed by These Characteristics on "Workflow Orchestration"
Solid tumor data's diverse sources and inconsistent update cycles require flexible data ingestion and version management mechanisms in the workflow. Unstructured and semi-structured document structures mean more resources are needed for information extraction and structuring during data preprocessing. For instance, accurately extracting critical fields like tumor size and lymph node metastasis status from pathology reports requires advanced Named Entity Recognition (NER) and relation extraction capabilities. The presence of specialized medical terminology and abbreviations challenges knowledge base construction and query understanding in the Retrieval-Augmented Generation (RAG) phase. This requires high-quality medical dictionaries and synonym lists. Additionally, the standardization of fields and units necessitates strict unit consistency checks during information comparison, validation, and report generation to prevent data errors from unit confusion. Multi-language documents require ensuring that text processing models in the workflow support multi-language capabilities or perform unified language conversion during preprocessing.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Chunk Size) | 800–1200 characters | Solid tumor document paragraphs are typically long and contain multiple key pieces of information. This range helps maintain semantic integrity and reduces context loss due to splitting. |
Recall count (Recall Count) | Top 5 | This ensures enough relevant context is recalled, covering multiple aspects of complex solid tumor conditions, while avoiding interference from irrelevant information. |
Similarity threshold (Similarity Threshold) | 0.75–0.85 | Solid tumor documents have high similarity in specialized terminology. This threshold helps filter out highly relevant document segments and excludes generalized information. |
Rerank result count (Reranked Return Count) | 3 | After reranking, a small number of the most relevant segments are selected, improving the accuracy and conciseness of the final generation. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Processing large clinical trial reports or genetic testing reports can be time-consuming for file parsing. Increasing the timeout prevents parsing failures. |
maxContext | 32000 tokens | Solid tumor medical records and guidelines often contain a large amount of information. A larger context window can accommodate longer inputs, reducing information truncation. |
Three Common Mistakes
- Workflow execution times out, with
TASK_TIMEOUTerror code in logs. This can happen when parsing or vectorizing large solid tumor clinical reports, and the default timeout is insufficient. - Key medical fields in generated reports are empty or incorrect. This can occur if the information extraction model fails to accurately identify all required fields in specific pathology report formats, leading to upstream data quality issues propagating downstream.
- Multiple workflow calls unexpectedly trigger unintended workflow logic. This can happen if trigger conditions or
Agentrouting logic between workflows are unclearly defined, causing incorrect matching to workflows for other disease areas under specific inputs.
How to Verify Configuration
- Select solid tumor documents with complex medical terminology and multi-paragraph structures. Test if the workflow accurately extracts all predefined key fields and verifies the units and formats of the extracted values.
- Use a series of queries covering different solid tumor types (e.g., lung cancer, breast cancer) and disease descriptions. Check if the RAG module returns highly relevant document segments that support generating accurate answers.
- Simulate inputs with varying data volumes and document lengths. Observe workflow execution times to ensure that parameters like
PARSE_FILE_TIMEOUT_SECONDSeffectively prevent timeout errors when processing large documents. - Create multiple parallel workflows for different solid tumor subtypes. Use test inputs to verify that each workflow is correctly identified and executed, without confusion or false triggers.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.