Document Parsing and Chunking for Pharmacovigilance Products

Core data in pharmacovigilance consists of various reports and manuals. These documents originate from Marketing Authorization Holders (MAHs) and

Data Characteristics in this Domain

Core data in pharmacovigilance consists of various reports and manuals. These documents originate from Marketing Authorization Holders (MAHs) and include Periodic Safety Update Reports (PSURs), Individual Case Safety Reports (ICSRs), drug inserts, and clinical trial protocols and reports. Data updates frequently, especially for post-market drug safety data, which updates regularly or irregularly based on regulatory requirements. Document structures typically include standardized sections: basic drug information, indications, dosage and administration, adverse reactions, contraindications, precautions, and pharmacology/toxicology. The adverse reactions section often appears as lists or tables, containing medical terms, disease codes (e.g., MedDRA), drug dosages, and frequency. These documents use highly specialized language, often containing many technical terms and abbreviations. Units include dosage (mg, g), frequency (times/day), and time (days, weeks, months).

Constraints Imposed by these Characteristics on Document Parsing and Chunking

The specialized and structured nature of pharmacovigilance documents imposes specific requirements on document parsing and chunking. First, documents contain numerous medical terms and abbreviations. If chunk granularity is too small, it can lead to loss of context, affecting semantic understanding and recall accuracy. Second, reports contain tabular data, especially adverse event frequency and severity information. Parsing must fully retain table structures and data associations to prevent data misalignment or loss. Third, frequent document updates require efficient processing of incremental data and support for version management. Furthermore, these documents are often lengthy and contain multiple heading levels. Chunking must balance section completeness with information density to ensure subsequent retrieval returns sufficiently informative snippets.

Configuration Settings

Configuration ItemSuggested ValueRationale
Chunk size (Chunk Length)800–1200 charactersBalances the completeness of medical terminology context with information density for retrieval, avoiding insufficient semantics from overly short chunks.
Chunk Overlap Length (Chunk Overlap Length)100–200 charactersEnsures critical information at chunk boundaries is not lost due to splitting, especially at junctions of technical terms and short sentences.
File Type Limitpdf, docx, txtPharmacovigilance documents are often published in PDF format, also accommodating Word documents and plain text reports.
PARSE_FILE_TIMEOUT_SECONDS600 secondsAddresses long parsing times for large PSURs or clinical reports, preventing parsing failures due to timeouts.
Recall count (Number of Retrieved Chunks)Top 5–7 chunksProvides sufficient contextual information for the model to analyze complex queries while ensuring relevance of retrieval results.
Similarity threshold (Similarity Threshold)Calibrate by actual measurementAdjusts using a test set to balance recall and precision, accounting for the semantic similarity characteristics of medical texts.

Three Common Pitfalls

  • Timeout errors occur when parsing large PDF files. This happens because the PARSE_FILE_TIMEOUT_SECONDS parameter is set too low, not allowing enough processing time.
  • Table content in uploaded PDF documents is misaligned or data is lost after parsing. This occurs because appropriate table parsing strategies are not enabled or configured, leading to structured data being incorrectly treated as plain text.
  • Retrieved document snippets lack sufficient context, preventing the model from accurately understanding the query intent. This usually means Chunk size (Chunk Length) is set too short, splitting important related information into different snippets.

How to Confirm Correct Configuration

  • Upload typical pharmacovigilance reports (e.g., PSUR, ICSR). Check if parsed document snippets in the knowledge base are complete and semantically coherent, especially verifying table content correctness.
  • Perform retrieval tests for queries containing specific medical terms or adverse events. Check if the Recall count (Number of Retrieved Chunks) meets information needs and if relevance is within an acceptable range.
  • Monitor backend logs for file parsing failures or timeouts. Adjust parameters like PARSE_FILE_TIMEOUT_SECONDS based on error codes.
  • Use FastGPT's knowledge base preview function. Randomly sample parsed document snippets and manually verify their content quality, including the reasonableness of chunk boundaries.

The values provided are common starting points and should be measured against the reader's own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.