Data Characteristics
R&D documents in hematologic oncology have specific characteristics. Data sources include clinical trial reports, pathology analysis reports, gene sequencing data, research papers on drug mechanisms, and regulatory submissions. These documents update frequently, especially during new drug development and clinical research phases. Data streams often iterate weekly or even daily. Document structures are complex. They contain extensive unstructured text, such as physician diagnostic descriptions and patient follow-up records. They also include semi-structured data, for example, tabular clinical indicators, gene mutation sites, and drug dosages. Common fields include ECOG score, FISH test results, bone marrow biopsy reports, and gene fusion types. Units involve copy number, mutation frequency (%), cell percentage, and treatment cycle. Abbreviations and industry-specific terminology are common.
Constraints from these Characteristics on Multiturn Conversation and Prompts
The high update frequency of hematologic oncology R&D documents requires knowledge bases to synchronize with the latest data promptly. This prevents the use of outdated information in multiturn conversations. The complex document structure and extensive unstructured content limit the precision of single-keyword matching recall. This necessitates more refined semantic understanding and contextual association capabilities. For example, key information in a bone marrow biopsy report diagnostic description might be scattered across multiple paragraphs. In a multiturn conversation, a user might progressively inquire about the prognosis of a specific gene fusion type under different treatment regimens. This requires the system to continuously track the conversation topic and integrate information from various documents. Furthermore, domain-specific abbreviations and terms, such as CR (complete remission) and MRD (minimal residual disease), demand prompt designs that effectively guide the model to identify and correctly interpret these specialized terms, avoiding ambiguity. Accurate understanding of numerical fields (e.g., mutation frequency) also requires prompt numerical parsing capabilities to support queries and comparisons based on numerical ranges.
Configuration Settings
| Configuration Item | Suggested Value | Rationale |
|---|---|---|
Chunk size (Chunk Size) | 500–800 characters | Balances semantic completeness with recall efficiency. Avoids overly long chunks that dilute key information or overly short chunks that lose context. |
Recall count (Recall Count) | Top 8–12 entries | Considering the complexity of queries and the dispersed nature of information in hematologic oncology, increasing the recall count improves relevance coverage. |
Similarity threshold (Similarity Threshold) | 0.75–0.85 | Balances recall breadth and precision. Prevents recalling too many irrelevant chunks while ensuring highly relevant content is not missed. |
Rerank result count (Reranked Return Count) | Top 5 entries | After initial recall, a reranking model further optimizes the order, ensuring the most relevant core information is presented first to the user. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Hematologic oncology report files can be large, requiring longer parsing times. This provides sufficient time to prevent parsing failures due to timeouts. |
maxContext | 3000–4000 tokens | Supports longer multiturn conversation histories and more complex queries, especially when tracing multiple clinical indicators or treatment pathways. |
Three Common Pitfalls
- Uploaded clinical trial report files remain in an "parsing" state for an extended period and eventually fail to parse. This might be due to file size or complexity exceeding the
PARSE_FILE_TIMEOUT_SECONDSlimit, or the document content includes special formats that the model struggles to process. - In a multiturn conversation, the user asks about the efficacy of a specific
gene mutation siteunder different drugs, but the system fails to maintain context, reinterpreting the question each time. This occurs if the prompt design does not effectively guide the model to maintain conversation state or lacks effective summarization of past dialogue. - Model output contains incorrect or inaccurate descriptions of
ECOG scoreortreatment cycle. This might be due to data conflicts in the knowledge base, or the prompt fails to explicitly instruct the model to validate numerical information when citing it, leading to model hallucination.
How to Confirm Proper Configuration
- Select ten typical hematologic oncology clinical trial reports and pathology reports. Upload them to the system and check their parsing status. Ensure all files parse successfully within an acceptable timeframe.
- Design a series of multiturn conversation scenarios covering different depths and breadths. For example, from
leukemia classificationgradually delving intoprognosisandtreatment plansfor aspecific fusion gene. Verify that the system accurately understands and maintains conversation context. - For key fields like
ECOG score,FISH test results, andmutation frequency, construct queries that include numerical ranges and units. Confirm that the system returns accurate information and correctly handles numerical comparisons.
Note: The values provided are common starting points and should be measured against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.