Data Characteristics in This Category
Hematologic oncology clinical trial data comes from various sources. These include clinical trial registries (e.g., ClinicalTrials.gov, WHO ICTRP), medical literature databases (e.g., PubMed, Embase), internal reports from pharmaceutical companies or research institutions, and regulatory agency approval documents. Data update frequencies vary; registry information might update weekly or monthly, while literature databases could update daily. Document structures differ. Trial protocols often exist as PDFs, containing detailed inclusion/exclusion criteria, study designs, treatment regimens, and evaluation metrics. Registry entries are more structured, including fields like NCT ID, Condition, Intervention, Eligibility Criteria, and Locations. A specific characteristic of hematologic oncology is the large number of rare disease trials. These trials have smaller data volumes and often involve complex biomarker screening, leading to highly detailed inclusion and exclusion criteria.
Constraints Imposed by These Characteristics on Citation and Traceability
The high specialization and diverse document types of hematologic oncology clinical trial data impose specific constraints on citation and traceability. First, parsing PDF trial protocols can lead to loss of formatting information. This affects the accuracy of subsequent text segmentation and citation references. Second, complex biomarker information within inclusion/exclusion criteria, such as BCR-ABL1 fusion gene status or FLT3-ITD mutations, requires precise identification and citation to ensure rigorous pre-screening logic. Due to varying data update frequencies, cited content must specify the data source and extraction time to mitigate information lag risks. Additionally, rare disease trials have limited data, potentially resulting in sparse recall. This necessitates more refined similarity matching strategies. Ensuring the completeness of citation context is crucial to avoid misjudgment due to decontextualization.
Configuration Guidelines
| Configuration Item | Suggested Value | Rationale for This Value |
|---|---|---|
Chunk size (Segment Length) | 800–1200 characters | Ensures complete inclusion/exclusion criteria or treatment regimen descriptions are captured, preventing truncation of critical information. |
Recall count (Recall Count) | 10–15 items | Accounts for potentially small data volumes in rare disease trials, increasing the hit rate. |
Similarity threshold (Similarity Threshold) | 0.75–0.85 | Balances recall precision and recall rate, reducing missed recalls due to differences in specialized terminology. |
Rerank result count (Reranked Return Count) | 5–7 items | Selects the most relevant citation snippets, avoiding interference from excessive irrelevant information. |
Citation Metadata Fields | source_url, document_title, update_date | Clearly identifies the source, document name, and data update time, facilitating user verification and traceability. |
Parsing Timeout | 600 seconds | Handles parsing of large PDF documents, especially trial protocols containing complex tables or figures. |
Three Common Pitfalls
- Knowledge base answers do not display citations or show "No permission to operate this conversation record." This can be due to incomplete
Citation Metadata Fieldsconfiguration or improper permission settings for the citation display component. - Newlines are expected in the answer but
\nstrings are displayed instead. This typically occurs when the text parser does not correctly handle newlines, or the output rendering layer does not interpret\nas a newline character. - The number of citations is too high or too low, affecting answer quality. This might be because
Recall count(Recall Count) orSimilarity threshold(Similarity Threshold) were not optimized for hematologic oncology data characteristics.
How to Confirm Proper Configuration
- Verify that each answer's citation source clearly points to the original document or registry entry, and is directly accessible via the
source_urlfield. - Check if citation snippets accurately include key information such as patient characteristics, disease staging, and biomarker status, to confirm the correctness of the pre-screening logic.
- Evaluate the combined effect of
Recall count(Recall Count) andSimilarity threshold(Similarity Threshold) by simulating various complex queries. This ensures that highly relevant and comprehensive citations are recalled.
The values provided are common starting points and should be measured against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.