Data Characteristics
Data for hematologic oncology clinical trial pre-screening primarily comes from medical literature, clinical trial protocols, patient medical records, genetic testing reports, and drug development materials. This data updates frequently; new clinical research, drug information, and treatment guidelines are regularly published. Document structures are complex. For instance, clinical trial protocols often include multi-level headings, tables, figures, appendices, and revision histories. Patient records contain unstructured physician notes, test results, and imaging reports.
Fields and units are highly specialized, covering various diagnostic metrics such as complete blood count, bone marrow cytology, flow cytometry, FISH, and gene sequencing. Units include g/L, ng/mL, %, and copy number, often accompanied by normal ranges or critical values. Additionally, much critical information appears in tables or nested lists. Images may also contain important morphological or pathological findings.
Constraints on Document Parsing and Chunking
The complexity of hematologic oncology pre-screening data imposes several constraints on document parsing and chunking. First, multiple data sources and high update frequency require the knowledge base to support efficient document synchronization and incremental updates for timely information. Second, complex document structures mean a single text segmentation strategy cannot effectively capture contextual relationships. For example, inclusion/exclusion criteria in clinical trial protocols may span different sections, requiring cross-paragraph or even cross-page correlation.
Key biomarkers, gene mutation information, and treatment plans contained in tables and images, if not effectively extracted, directly impact pre-screening accuracy. Finally, specialized fields, units, and a large volume of medical terminology demand high domain adaptability from tokenizers and entity recognition models. This prevents loss or misinterpretation of critical information due to inaccurate lexical analysis.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
UPLOAD_FILE_MAX_SIZE | 500 MB | Clinical trial protocols and medical records often contain numerous images and tables, leading to large file sizes. |
Chunk size (Chunk Length) | 800–1200 characters (characters) | Balances contextual completeness with recall efficiency, preventing semantic fragmentation from excessive splitting or interference from overly long, irrelevant information. |
Maximum Paragraph Depth (Max Paragraph Depth) | 5 | Clinical trial protocols typically have deeply nested chapter structures; this ensures parsing of critical sub-section content. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds (seconds) | Processing large PDF files (e.g., those with high-resolution pathology images) can be time-consuming; this prevents parsing timeouts. |
Enable OCR | Yes | Patient medical records and some older literature often exist as scanned documents, requiring text recognition from images. |
Recall count (Recall Count) | 10–15 entries (items) | Hematologic oncology pre-screening decisions are complex, requiring more supporting evidence to improve recall rate. |
Common Pitfalls
- Uploaded PDF files fail to parse text content, returning empty knowledge chunks. This occurs when PDF files are pure image scans and OCR is not enabled or the OCR service fails to recognize text.
- Descriptions of a specific gene mutation in the knowledge base are incomplete, lacking critical contextual information. This likely happens when the chunk length is set too short, causing critical information to be split across different knowledge chunks and losing association.
- Clinical guidelines containing many tables are imported, but queries cannot effectively retrieve data from these tables. This may be because the document parser does not specifically handle table structures, treating table content as plain text for segmentation.
How to Verify Configuration
- Upload representative documents containing complex tables, images, and multi-level headings. Check if the parsed knowledge chunks include all critical information and verify that table data is correctly extracted.
- Construct query statements using medical terms and gene mutation names specific to hematologic oncology. Check if the knowledge base recalls accurate knowledge chunks containing these terms and evaluate the completeness of the recalled chunks.
- Parse documents from various sources (e.g., clinical trial protocols, patient medical records, genetic testing reports). Observe if parsing timeouts or errors occur, and check log outputs to determine if the parser is functioning correctly.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.