Data Characteristics
Data for hematologic oncology clinical trial pre-screening primarily comes from global clinical trial registries (e.g., ClinicalTrials.gov, WHO ICTRP), trial protocols published by pharmaceutical companies and research institutions, medical journal literature, and anonymized patient medical records. This data updates frequently; monthly or even weekly updates are common, especially for new drug development and trial progress. Document structures are diverse, including structured trial registration forms, semi-structured PDF research protocols, and unstructured medical report texts. Fields and units are highly specialized, such as "Complete Remission Rate (CR)," "Progression-Free Survival (PFS)," and "Minimal Residual Disease (MRD)," often accompanied by specific detection methods and evaluation criteria. Gene mutation information (e.g., FLT3-ITD, IDH1/2) and chromosomal abnormalities (e.g., t(15;17)) are important features; their naming and reporting formats are relatively fixed.
Constraints Imposed on Knowledge Base Retrieval and Recall by These Characteristics
The high update frequency of hematologic oncology data requires the knowledge base to support rapid synchronization and incremental updates. This prevents the retrieval of outdated or inaccurate trial information. Diverse document structures necessitate flexible text parsing strategies, for example, extracting key paragraphs from PDFs or identifying specific fields from structured data. Highly specialized fields and units require embedding models to accurately understand the semantics of medical terms and distinguish subtle differences, such as the impact of different gene mutation sites on drug sensitivity. For specific identifiers like gene mutations and chromosomal abnormalities, exact matching is more critical than semantic similarity. This requires effective handling of abbreviations, aliases, and standard nomenclature during recall. Additionally, the complexity of clinical trial protocols often results in long individual documents, requiring fine-grained segmentation strategies to ensure an appropriate retrieval granularity.
Configuration Settings
| Configuration Item | Suggested Value | Rationale |
|---|---|---|
Chunk size (Segment Length) | 500–800 characters | Clinical trial protocols contain a large amount of information per document. This length maintains contextual completeness and prevents truncation of key information. |
Chunk Overlap Rate (Segment Overlap Rate) | 10%–15% | Appropriate overlap helps capture semantic connections across paragraphs, especially when describing trial objectives and inclusion/exclusion criteria. |
Recall count (Number of Retrieved Items) | Top 10–15 | The complexity of hematologic oncology trials requires a higher recall volume to cover potentially relevant trials. |
Similarity threshold (Similarity Threshold) | 0.75–0.85 | This ensures retrieved results are highly relevant to the query, filtering out a large number of irrelevant or generic trial descriptions. |
Rerank result count (Number of Reranked Items) | Top 5 | After retrieving a higher number of items, a reranking model further selects the most relevant trials, improving final precision. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Processing large PDF format trial protocols can be time-consuming. This value helps prevent parsing timeouts. |
Common Mistakes
- After a knowledge base update, search results do not reflect the latest trial information. This happens because the incremental update mechanism is not correctly configured or the scheduled task does not trigger.
- Retrieving specific gene mutations (e.g.,
BCR-ABL1) yields inaccurate results. This occurs when specialized terms are not enhanced with a vocabulary or synonym mapping. - Uploading large clinical trial protocol PDF files results in an
embedding error. This happens when the file size exceeds theUPLOAD_FILE_MAX_SIZElimit or parsing times out.
How to Verify Configuration
- For multiple recently updated hematologic oncology clinical trials, execute queries and verify whether the recall results include these latest trials and whether relevant fields are accurate.
- Select a set of queries containing specific gene mutations or chromosomal abnormalities. Check whether the recall results precisely match trials containing these identifiers.
- Upload a typical long clinical trial protocol PDF document. Observe whether its segment count and content meet expectations, and whether the embedding process completes normally.
Note: The values given are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.