Data Characteristics
Attenuated inactivated vaccine pharmacovigilance data originates primarily from clinical trial reports, real-world evidence (RWE) reports, post-market surveillance reports, and global adverse event databases. These documents are typically in PDF format and contain a mix of structured and unstructured text. Structured sections may include tables detailing adverse event incidence rates, patient demographic information, and laboratory test results. Unstructured sections cover detailed adverse event descriptions, medical assessments, and follow-up records. Data update frequency is influenced by regulatory requirements and the vaccine lifecycle; for example, post-market surveillance reports may be submitted quarterly or annually. Documents frequently contain medical terminology, drug batch information, dosage units (e.g., IU, μg), and event timestamps.
Constraints on Document Parsing and Chunking
The complexity of attenuated inactivated vaccine pharmacovigilance documents imposes specific requirements on document parsing and chunking. First, the mixed layout of images and text in PDF documents necessitates that parsers can process text content within images, especially adverse event reports from scanned documents. Second, data within structured tables, such as adverse event codes (e.g., MedDRA terms) and frequencies, requires accurate extraction while preserving contextual relevance to avoid data silos. The specialized and diverse nature of medical terminology means that chunking based solely on characters or punctuation marks can disrupt critical information units. Moreover, time-series descriptions of adverse events and causal relationship analyses depend on integrating information across paragraphs and even pages. Therefore, chunking strategies must balance fine-grained information extraction with macroscopic contextual understanding.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size | 500–800 characters | Balances the completeness of adverse event descriptions with the granularity of recall, avoiding the truncation of critical medical information. |
Chunk overlap | 100–150 characters | Ensures contextual continuity, especially when describing adverse reactions or medical assessments across paragraphs. |
Knowledge Base Type | Text Knowledge Base | Suitable for processing large volumes of unstructured text content and supports text embedding and vector retrieval. |
Parsing Mode | Smart Chunking | Prioritizes using the model to understand semantic boundaries, handling complex paragraph structures in medical reports. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Addresses the parsing time requirements for large clinical trial reports or annual summary reports. |
UPLOAD_FILE_MAX_SIZE | 500 MB | Allows uploading PDF documents containing numerous charts and scanned pages. |
Common Pitfalls
- Document parsing results in extensive blank spaces or garbled text. This occurs when image text recognition capabilities are not enabled or configured, leading to the loss of critical adverse event information from scanned documents.
- Retrieved chunks lack essential context, for example, only recalling an adverse event name without dosage or medication history. This happens when the chunk length is too short, causing semantic units to be improperly split.
- Uploading large PDF documents results in timeout or failure messages. This is because the file parsing timeout parameter
PARSE_FILE_TIMEOUT_SECONDSis set too low, failing to accommodate the processing time for complex documents.
Verification Steps
- Select typical attenuated inactivated vaccine clinical reports, upload them, and inspect the parsed text content. Confirm that key medical terms, adverse event descriptions, and table data are complete and free of garbling.
- Perform retrieval tests on the parsed knowledge base using queries that include medical terminology and specific symptoms. Check if the retrieved results contain relevant contextual information and evaluate their relevance.
- Monitor backend logs to confirm that file parsing tasks complete normally without timeout or parsing failure error codes when processing documents of varying sizes and complexities.
The values provided are common starting points and should be measured against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.