Document Parsing and Chunking for Phase II-III Clinical Pharmacovigilance

Phase II-III clinical trial pharmacovigilance data originates from clinical trial protocols, case report forms (CRFs), serious adverse event (SAE)

Data Characteristics

Phase II-III clinical trial pharmacovigilance data originates from clinical trial protocols, case report forms (CRFs), serious adverse event (SAE) reports, periodic safety update reports (PSUR/DSUR), and investigator brochures. These documents are typically in PDF format. They contain structured and unstructured text, including patient demographics, medical history, adverse event descriptions, laboratory results, diagnostic codes (e.g., MedDRA coding), and assessments of adverse event-drug causality. Data updates frequently during trials, especially for SAE reports, which require submission within specific deadlines. Document structures are complex, often including tables, figures, and lengthy narrative paragraphs. Fields and units are highly specialized, for example, dose units (mg/kg), time units (days, hours), laboratory indicator units (mmol/L, U/L), and medical terminology.

Constraints on Document Parsing and Chunking

The complexity of Phase II-III clinical pharmacovigilance data imposes specific requirements on document parsing and chunking. First, accurate extraction of table and figure content from PDF documents requires advanced layout recognition capabilities from the parser. Second, the real-time nature and high frequency of SAE report updates mean the knowledge base must support incremental updates and version management to avoid duplicate ingestion or information loss. Key information in lengthy narrative paragraphs, such as the progression of adverse events, severity judgments, and management measures, must be effectively identified and chunked. The specialized nature of medical terminology and diagnostic codes requires chunking strategies to maintain the integrity of this context, preventing semantic loss due to fragmentation. For example, an adverse event description may span multiple sentences or even paragraphs. Chunking must ensure related information is grouped into the same chunk to support precise retrieval and question answering.

Configuration Settings

Configuration ItemRecommended ValueRationale
Chunk size800–1200 charactersBalances long text context and retrieval efficiency, avoids cutting off core information.
Overlap Length100–200 charactersEnsures contextual continuity at chunk boundaries, improves recall.
PARSE_FILE_TIMEOUT_SECONDS600 secondsAccommodates parsing large PDF documents, prevents timeout errors.
maxContextCalibrate by measurementBased on actual LLM context window limits, ensures effective input.
Recall countTop 5–10 itemsBalances retrieval accuracy and LLM processing load.
Similarity threshold0.75–0.85Filters irrelevant chunks, improves retrieval result quality.

Common Pitfalls

  • Incomplete retrieval results from Excel data in the knowledge base. For example, querying "usage location is abroad" returns too few entries. This often occurs when Excel files are not correctly parsed to identify all data rows or columns, leading to some data not being effectively indexed.
  • System errors when uploading large PDF documents, indicating timeouts or processing failures. This may relate to an insufficient PARSE_FILE_TIMEOUT_SECONDS parameter setting, which cannot handle the time required for complex document parsing.
  • Auxiliary data, such as MedDRA codes, do not contribute to matching within chunked content, leading to imprecise retrieval results. This indicates the chunking strategy may not have effectively associated or indexed auxiliary data with the primary content.

Verification Steps

  • Upload typical Phase II-III clinical trial protocols or SAE reports. Examine the parsed knowledge base chunks to ensure key information, such as patient characteristics, adverse event descriptions, and dosage information, is complete and contextually coherent.
  • Perform retrieval using queries containing specific medical terms or diagnostic codes. Verify that the recalled chunks accurately contain this specialized information and that relevant auxiliary data is effectively utilized.
  • Conduct random sample checks of parsed chunks. Ensure that table content and key data points from lengthy narrative paragraphs (e.g., laboratory indicator values, time points) are correctly extracted and chunked, and that no semantic fragmentation occurs.

The values provided are common starting points and should be measured against your own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.