Document Parsing and Chunking for mRNA Vaccine Pharmacovigilance

mRNA vaccine pharmacovigilance data comes from clinical trial reports, real-world evidence (RWE) studies, case reports, regulatory post-market

Data Characteristics

mRNA vaccine pharmacovigilance data comes from clinical trial reports, real-world evidence (RWE) studies, case reports, regulatory post-market surveillance reports, and peer-reviewed journal articles. These documents are typically in PDF, DOCX, or XML formats. Data updates frequently, especially post-market surveillance data, which may update weekly or monthly. Document structures are complex. They include unstructured text (e.g., adverse event descriptions), semi-structured tables (e.g., adverse reaction lists, dosage information), and structured fields (e.g., batch numbers, report IDs, patient IDs). Fields and units follow medical and pharmaceutical standards. Examples include MedDRA terms for adverse event coding, WHO-ART terms, and ICD-10 codes. Common dosage units are milligrams (mg), micrograms (µg), and milliliters (mL). Time units are hours, days, and weeks.

Constraints on Document Parsing and Chunking

The complexity of mRNA vaccine pharmacovigilance documents imposes specific requirements on document parsing and chunking. High update frequency requires incremental parsing and rapid knowledge base updates. Complex document structures demand parsers that handle plain text and effectively identify and extract table data, distinguishing different semantic regions. For example, adverse event descriptions are typically long texts, while adverse reaction lists are structured tables. Standardized medical terminology and units require parsers with domain knowledge to ensure accurate identification and segmentation, preventing information loss or misunderstanding due to ambiguous terms. Large files, such as clinical trial reports, are common. This requires high performance and stability in the parsing process to avoid timeouts. Chunking strategies must consider the completeness of adverse event reports, preventing critical information from being split across different chunks.

Configuration Settings

Configuration ItemRecommended ValueRationale
UPLOAD_FILE_MAX_SIZE500 MBClinical trial reports and large research documents can reach tens or even hundreds of megabytes. This ensures successful uploads.
Chunk size (Chunk Length)800–1200 characters (characters)Balances context completeness with recall accuracy, accommodating the long text characteristics of adverse event descriptions.
Chunk Overlap Length (Chunk Overlap Length)100–150 characters (characters)Ensures critical medical terms and phrases are not truncated at chunk boundaries, maintaining semantic coherence.
PARSE_FILE_TIMEOUT_SECONDS600 seconds (seconds)Provides sufficient parsing time for large PDF/DOCX files, preventing timeout interruptions.
Custom Separator (Custom Separators)。!?\n\n###Segments based on common periods, question marks, exclamation marks, paragraph breaks, and specific section titles in adverse event reports.
Table Recognition ModeSmart RecognitionEffectively extracts semi-structured data like adverse reaction lists and dosage tables found in clinical trial reports.

Common Pitfalls

  • Parsing large PDF documents results in timeout errors, preventing successful file import. This occurs when the PARSE_FILE_TIMEOUT_SECONDS configuration is too low for complex document parsing times.
  • Imported Excel table data appears as multiple rows merged into a single chunk in the knowledge base. This happens when Custom Separator (Custom Separators) fails to correctly identify row delimiters in tables, or when table recognition is not enabled.
  • Key terms in adverse event descriptions are not effectively matched during recall. This is because Chunk size (Chunk Length) is too long or Chunk Overlap Length (Chunk Overlap Length) is insufficient, diluting important context or splitting critical information.

How to Verify Configuration

  • Upload multiple typical large PDF clinical trial reports. Check if the file parsing status shows "success" and no timeout errors.
  • Import DOCX or Excel files containing tables. Check the chunk content in the knowledge base to confirm that table data is correctly identified and chunked by row or logical unit.
  • Perform knowledge base retrieval tests for specific medical terms and phrases in adverse event descriptions. Observe whether the recall results include the expected relevant chunks and check the completeness of the chunk content.

The values provided are common starting points and should be measured against your own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.