Data Characteristics
Hemato-oncology pharmacovigilance data originates from clinical trial reports, real-world evidence (RWE) data, individual case safety reports (ICSRs), medical literature, and regulatory guidelines and alerts. These documents update frequently, especially with new clinical trial results and regulatory safety updates. Document structures vary, including structured tabular data (e.g., patient characteristics, adverse event codes), semi-structured narrative text (e.g., clinician descriptions of adverse events, case summaries), and unstructured free text. Fields include general demographic information, disease staging, gene mutation types, treatment regimens, drug dosages, and CTCAE (Common Terminology Criteria for Adverse Events) grades for AEs (adverse events) and SAEs (serious adverse events), and drug causality assessments. Units involve dosage (mg, g), time (days, weeks, months), and biomarker values.
Constraints from "Document Parsing and Chunking"
The diversity of hemato-oncology documents requires highly robust document parsing. Precise extraction of drug dosages and gene mutation information from structured tables is critical; any deviation can affect subsequent pharmacovigilance analysis. In semi-structured and unstructured text, clinician descriptions of adverse events often contain specialized terminology, abbreviations, and context-dependent expressions. Chunking must effectively identify and retain this critical information, preventing loss of context due to excessive splitting. Frequent updates, particularly new medical literature and regulatory alerts, mean the parsing process must respond efficiently to ensure knowledge base timeliness. The presence of specialized coding systems like CTCAE grades requires the parser to handle multi-level classification information and associate it with free-text descriptions.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
UPLOAD_FILE_MAX_SIZE | 500 MB | Clinical trial reports or research literature in hemato-oncology often contain numerous charts and detailed data, resulting in large file sizes. A sufficiently large upload limit is necessary. |
Chunk size (Chunk Length) | 800–1200 characters (characters) | Balances the completeness of adverse event descriptions with retrieval efficiency, preventing the splitting of critical contextual information. |
Overlap Length | 100 characters (characters) | Ensures semantic continuity at chunk boundaries, especially when processing clinical narrative text, preventing important information from being fragmented. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds (seconds) | Parsing large files, particularly PDFs with complex tables and multi-level text, requires more processing time to avoid timeouts. |
Enable PDF High-Precision Parsing | Enabled | Ensures accurate identification and extraction of tables, captions, and complex layouts common in medical literature within PDFs. |
Text Cleaning Rules | Remove headers/footers, consecutive whitespace | Reduces noise, improving the accuracy of subsequent vectorization and retrieval, while retaining core medical information. |
Common Pitfalls
- After uploading a PDF file, an
OCR Errorappears, and logs show image processing failure. This can occur if the PDF scan quality is poor or if it contains non-text images that the OCR engine cannot recognize. - Key drug dosage or adverse event grade fields are empty in the parsed file content. This typically happens when document structure changes or the parsing template does not adapt to new table formats, leading to failed structured information extraction.
- When processing large clinical trial reports, FastGPT shows a
Gateway Timeouterror. This indicates that the file parsing time exceeded the default gateway or service timeout limit.
Verification Steps
- Upload a typical hemato-oncology clinical trial report PDF. Check if the parsed text content completely retains drug names, dosages, adverse event descriptions, and CTCAE grades.
- Randomly select parsed document segments and compare them against the original document. Verify if chunk boundaries are reasonable and if any important medical terminology or context is improperly split.
- Use documents containing complex tables for parsing. Verify if data within tables (e.g., patient baseline characteristics, laboratory indicators) is correctly identified and converted into a usable text format.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.