Document Parsing and Chunking for ADC Pharmacovigilance

Antibody-Drug Conjugate (ADC) pharmacovigilance data originates from clinical trial reports, real-world evidence (RWE) studies, post-market

Data Characteristics

Antibody-Drug Conjugate (ADC) pharmacovigilance data originates from clinical trial reports, real-world evidence (RWE) studies, post-market surveillance reports, spontaneous adverse event reports, and scientific literature. This data updates frequently, especially post-market surveillance reports, which are released periodically according to regulatory requirements. Document structures typically include detailed patient information, medical history, adverse event descriptions, severity, outcomes, causality assessments, and drug batch information. Fields often cover patient demographics, disease diagnoses, ADC drug generic and brand names, dosage, administration routes, frequency, MedDRA codes for adverse events, CTCAE grades, and laboratory test results (e.g., liver and kidney function, blood counts). Units encompass common dosage units (mg/kg), time units (days, weeks, months), and international or conventional units for various laboratory indicators.

Constraints from these Characteristics on Document Parsing and Chunking

The multi-source nature and high update frequency of ADC pharmacovigilance data require the document parsing system to have efficient automated processing capabilities to handle continuous influxes of new data. Multiple nested levels within document structures (e.g., symptoms, signs, and diagnoses under adverse event descriptions) challenge parsing accuracy, requiring the ability to identify and extract deep semantic information. ADC-specific targets, linkers, and cytotoxic components mean adverse event descriptions may contain highly specialized biomedical terminology and abbreviations. This is crucial for accurate tokenization and entity recognition. Furthermore, clinical trial reports often include tables and figures, requiring the parser to handle non-text content and convert it into retrievable structured information. Ambiguous descriptions or non-standard expressions in adverse event reports also demand chunking strategies that maintain contextual integrity, preventing critical information from being fragmented.

Configuration Guidelines

Configuration ItemRecommended ValueRationale
Chunk size800–1200 charactersRetains the complete context of adverse event descriptions while preventing individual chunks from becoming too long and diluting key information.
Chunk Overlap Length100–200 charactersEnsures critical information spanning across paragraphs is not lost due to boundary cuts, improving recall.
Parsing ModeSmart ChunkingAdapts to the diverse structures in ADC pharmacovigilance documents, including paragraphs, lists, and tables.
PARSE_FILE_TIMEOUT_SECONDS600 secondsProvides sufficient parsing time for large clinical trial reports or PDF documents.
Max File Size200 MBSupports uploading clinical study reports or summary files containing large amounts of raw data.
Entity Recognition ModelMedical Pre-trained ModelAccurately identifies specialized entities such as ADC drug names, adverse event terms, and laboratory indicators.

Common Pitfalls

  • Uploaded PDF documents show parsing failure or content loss. This usually occurs because the PDF contains too many scanned images or non-text layers, leading to inaccurate OCR recognition or the parser's inability to extract text.
  • Poor or irrelevant recall results when searching for specific adverse events after import. This may be due to a short chunk length, which severs critical causal descriptions or drug-event association information, leading to context loss.
  • Remaining in an "indexing" state for an extended period without successful ingestion. This could be related to a low PARSE_FILE_TIMEOUT_SECONDS parameter, where large or complex documents fail to complete parsing within the allotted time.

Configuration Validation

  • Select a PDF document containing typical ADC pharmacovigilance adverse event descriptions. Upload it and check if the parsed text content is complete and free of garbled characters. Pay particular attention to MedDRA codes and CTCAE grades for adverse events.
  • Choose snippets from the document that involve drug names, dosages, and adverse event descriptions. Perform a recall test using the search function. Evaluate the accuracy and relevance of the returned results. Adjust Chunk size and Chunk Overlap Length until satisfactory.
  • Upload a clinical trial report containing tables and figures. Check if table content is correctly parsed and converted into retrievable text. Verify if text descriptions next to figures are effectively extracted.

The values provided are common starting points and should be measured against the reader's own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.