Data Characteristics
Cardiovascular pharmacovigilance data comes from various sources. These include clinical trial reports, real-world evidence (RWE) databases, individual case safety reports (ICSRs), and medical literature. Data update frequencies vary. Clinical trial reports are typically released after study completion. ICSRs and RWE databases may update in real-time. Document structures also differ. Clinical trial reports are often structured PDFs with sections like introduction, methods, results, and discussion. ICSRs are usually semi-structured XML or PDFs, containing fields for patient information, drug details, and adverse event descriptions. Specific cardiovascular events, such as myocardial infarction or arrhythmia, are key fields. Units involve physiological indicators like blood pressure (mmHg), heart rate (bpm), and electrocardiogram (e.g., QT interval, ms). Precise identification of these is necessary.
Constraints Imposed by These Characteristics on Document Parsing and Chunking
The heterogeneous nature of cardiovascular pharmacovigilance data requires document parsers to handle multiple file formats. Structured PDFs and semi-structured XML are particularly important. Real-time updates demand efficient parsing for rapid data ingestion and processing. Diverse document structures necessitate flexible chunking strategies. For instance, clinical trial reports may require chunking by section to maintain contextual integrity. ICSRs might need fine-grained chunking by adverse event or drug information fields. Identifying specific cardiovascular fields and units is critical. This ensures "myocardial infarction" is not misidentified as "myocardial strain." It also ensures correct extraction of values and units, distinguishing 120/80 mmHg from 80 bpm. This directly impacts subsequent entity extraction and relationship identification. Image content, such as ECG waveforms or imaging screenshots, requires additional processing mechanisms if they contain critical information.
Configuration Settings
| Configuration Item | Suggested Value | Rationale |
|---|---|---|
Chunk size | 800–1200 characters | Balances the completeness of cardiovascular event descriptions with retrieval granularity, avoiding truncation of key medical terms. |
Chunk overlap | 100–200 characters | Ensures contextual continuity at chunk boundaries, especially when describing complex cardiovascular adverse events. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Accommodates the parsing time for large clinical trial reports or PDFs with extensive charts. |
UPLOAD_FILE_MAX_SIZE | 1000 MB | Supports uploading PDF files containing high-resolution images or many pages. |
EnabledImage Recognition | Yes | Identifies and extracts critical visual information from ECGs and imaging reports, such as myocardial ischemia areas. |
PDF Parsing Mode | Structured Parsing Priority | Prioritizes using the PDF's internal structural information (e.g., table of contents, headings) to improve parsing accuracy. |
Common Pitfalls
- Uploaded PDF files have missing content, particularly image information that was not recognized. This occurs if the default parser does not have image OCR enabled or if image quality is too low for OCR to succeed.
- HTML interface documentation is uploaded, but the knowledge base fails to parse effective content. This happens because HTML documents have complex structures, often containing code or tables, and default text extraction strategies cannot effectively identify key information.
- Parsed chunks have incoherent context, leading to inaccurate retrieval results. This occurs if
Chunk sizeis set too small, truncating key sentences describing complete cardiovascular events, or ifChunk overlapis not configured.
How to Verify Configuration
- Upload a PDF report containing ECGs or angiograms. Check if the text information within the images can be retrieved from the knowledge base.
- Upload a typical ICSR report (XML or PDF). Check if the parsed chunks correctly differentiate fields such as patient basic information, drug details, and adverse event descriptions.
- Randomly select parsed document chunks. Manually evaluate their semantic completeness. Ensure cardiovascular medical terms and numerical units are not truncated.
- Perform simulated queries. Input long-tail questions related to cardiovascular drug adverse reactions. Check if retrieval results are accurate and include necessary contextual information.
The values provided are common starting points. Measure them against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.