Document Parsing and Chunking for Lead Optimization in Pharmacovigilance

Pharmacovigilance data during lead optimization primarily comes from preclinical research reports, toxicology reports, pharmacokinetic reports, early

Data Characteristics

Pharmacovigilance data during lead optimization primarily comes from preclinical research reports, toxicology reports, pharmacokinetic reports, early clinical trial (e.g., Phase I) data, and some in vitro and in vivo pharmacodynamic study documents. These documents are often in PDF format, with some in Word or structured data tables. The update frequency is relatively low, typically aligning with experimental progress or submission of phased reports. Document structures often follow industry standards or internal templates, including sections like abstracts, materials and methods, results, and discussion. Fields and units are highly specialized, such as dose (mg/kg), concentration (µM), half-life (h), and toxicity indicators (e.g., AST, ALT, LD50), often accompanied by complex charts and statistical data.

Constraints Imposed by These Characteristics on Document Parsing and Chunking

Diverse document sources require parsing tools to be compatible with multiple file formats. The specialized structure and terminology of reports mean traditional, general-purpose chunking methods may struggle to accurately identify key information boundaries. For example, in a toxicology report, results from different toxicity tests might be scattered across several sub-sections, requiring fine-grained chunking to maintain semantic integrity. Complex charts and embedded images, if not effectively parsed and converted into a format understandable by large language models (LLMs), will lead to information loss. Furthermore, specific fields and units, such as dose-response curve data, must not be incorrectly truncated or lose context during chunking, which would affect the LLM's accuracy in identifying drug safety signals. Low document update frequency means parsing accuracy is critical to reduce repeated parsing costs.

Configuration Settings

Configuration ItemRecommended ValueRationale
Chunk Length800–1200 charactersBalances semantic completeness with model context window limits, adapting to paragraph lengths in specialized reports.
Overlap Length100–150 charactersEnsures contextual continuity between chunks, preventing critical information from being cut off.
File TypesPDF, DOCX, XLSXCovers common report formats in the lead optimization phase.
Image ParsingEnabledEnsures key data in charts and embedded images can be extracted or described.
PARSE_FILE_TIMEOUT_SECONDS600 secondsAccommodates the parsing time for large toxicology reports or documents with complex charts.
Text Cleaning RulesCalibrated by actual measurementPre-processes specialized terminology and symbols to improve parsing accuracy.

Three Common Mistakes

  • Missing specialized terms or key numerical values in parsing results. This occurs when chunk length is set too short, causing critical information to be split across different chunks, or when text cleaning rules are too aggressive.
  • Table data in PDF documents is not correctly identified, leading to table content being parsed as unstructured text. This happens when the PDF parser's ability to recognize complex table structures is insufficient, or when the table structure preservation option is not enabled.
  • Some document parsing takes too long, or even results in timeout errors. This happens when the PARSE_FILE_TIMEOUT_SECONDS parameter is set too low, failing to account for the parsing demands of large or highly complex documents.

How to Verify Configuration

  • Select a typical report containing various data types (text, tables, images) for parsing. Check if the parsed chunks are complete and without critical information omissions.
  • Randomly select parsed chunks and verify the accuracy of specialized terms, dosage units, and experimental results against the original document.
  • Test the parsing efficiency of documents of different sizes and complexities. Record parsing times and compare them against the PARSE_FILE_TIMEOUT_SECONDS parameter to confirm no timeouts occur.

The values provided are common starting points. Always measure against your own samples to determine the optimal configuration for specific use cases.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.