Document Parsing and Chunking for Cardiovascular Pharmacovigilance

Cardiovascular pharmacovigilance data comes from various sources. These include clinical trial reports, real-world evidence (RWE) databases

Data Characteristics

Cardiovascular pharmacovigilance data comes from various sources. These include clinical trial reports, real-world evidence (RWE) databases, individual case safety reports (ICSRs), and medical literature. Data update frequencies vary. Clinical trial reports are typically released after study completion. ICSRs and RWE databases may update in real-time. Document structures also differ. Clinical trial reports are often structured PDFs with sections like introduction, methods, results, and discussion. ICSRs are usually semi-structured XML or PDFs, containing fields for patient information, drug details, and adverse event descriptions. Specific cardiovascular events, such as myocardial infarction or arrhythmia, are key fields. Units involve physiological indicators like blood pressure (mmHg), heart rate (bpm), and electrocardiogram (e.g., QT interval, ms). Precise identification of these is necessary.

Constraints Imposed by These Characteristics on Document Parsing and Chunking

The heterogeneous nature of cardiovascular pharmacovigilance data requires document parsers to handle multiple file formats. Structured PDFs and semi-structured XML are particularly important. Real-time updates demand efficient parsing for rapid data ingestion and processing. Diverse document structures necessitate flexible chunking strategies. For instance, clinical trial reports may require chunking by section to maintain contextual integrity. ICSRs might need fine-grained chunking by adverse event or drug information fields. Identifying specific cardiovascular fields and units is critical. This ensures "myocardial infarction" is not misidentified as "myocardial strain." It also ensures correct extraction of values and units, distinguishing 120/80 mmHg from 80 bpm. This directly impacts subsequent entity extraction and relationship identification. Image content, such as ECG waveforms or imaging screenshots, requires additional processing mechanisms if they contain critical information.

Configuration Settings

Configuration ItemSuggested ValueRationale
Chunk size800–1200 charactersBalances the completeness of cardiovascular event descriptions with retrieval granularity, avoiding truncation of key medical terms.
Chunk overlap100–200 charactersEnsures contextual continuity at chunk boundaries, especially when describing complex cardiovascular adverse events.
PARSE_FILE_TIMEOUT_SECONDS600 secondsAccommodates the parsing time for large clinical trial reports or PDFs with extensive charts.
UPLOAD_FILE_MAX_SIZE1000 MBSupports uploading PDF files containing high-resolution images or many pages.
EnabledImage RecognitionYesIdentifies and extracts critical visual information from ECGs and imaging reports, such as myocardial ischemia areas.
PDF Parsing ModeStructured Parsing PriorityPrioritizes using the PDF's internal structural information (e.g., table of contents, headings) to improve parsing accuracy.

Common Pitfalls

  • Uploaded PDF files have missing content, particularly image information that was not recognized. This occurs if the default parser does not have image OCR enabled or if image quality is too low for OCR to succeed.
  • HTML interface documentation is uploaded, but the knowledge base fails to parse effective content. This happens because HTML documents have complex structures, often containing code or tables, and default text extraction strategies cannot effectively identify key information.
  • Parsed chunks have incoherent context, leading to inaccurate retrieval results. This occurs if Chunk size is set too small, truncating key sentences describing complete cardiovascular events, or if Chunk overlap is not configured.

How to Verify Configuration

  • Upload a PDF report containing ECGs or angiograms. Check if the text information within the images can be retrieved from the knowledge base.
  • Upload a typical ICSR report (XML or PDF). Check if the parsed chunks correctly differentiate fields such as patient basic information, drug details, and adverse event descriptions.
  • Randomly select parsed document chunks. Manually evaluate their semantic completeness. Ensure cardiovascular medical terms and numerical units are not truncated.
  • Perform simulated queries. Input long-tail questions related to cardiovascular drug adverse reactions. Check if retrieval results are accurate and include necessary contextual information.

The values provided are common starting points. Measure them against your own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.