Data Characteristics in This Domain
Pharmacovigilance data originates primarily from post-market surveillance reports, adverse event reports (ADRs), safety update reports (PSUR/DSUR), risk management plans (RMPs), and related regulatory documents. Document update frequency depends on the drug's lifecycle stage and regulatory requirements; for example, PSURs typically update semi-annually or annually. Document structures vary, including structured tables, semi-structured text, and free-text descriptions. Typical fields include drug name, batch number, dosage, administration route, adverse event terms (using MedDRA coding), event onset time, severity, outcome, and causality assessment. Units commonly involve time (days, weeks, months), dosage (mg, g, IU), and frequency (times/day).
Constraints Imposed by These Characteristics on Document Parsing and Chunking
Pharmacovigilance document characteristics impose specific requirements on document parsing and chunking. First, precise extraction of key identifiers like drug names and batch numbers is critical for data association accuracy. Second, adverse event descriptions often contain medical terminology and clinical details, demanding high semantic integrity in chunking to avoid context loss from over-chunking. Reports with high update frequencies require parsing processes to support incremental updates and version management. The mix of structured tables and free text in documents means a single parsing strategy is insufficient, necessitating a combination of table recognition and natural language processing techniques. Furthermore, the presence of specialized terms like MedDRA codes requires the parser to recognize and preserve their integrity, preventing truncation or misinterpretation during chunking, which would affect subsequent retrieval and analysis.
Configuration Settings
| Configuration Item | Recommended Value | Rationale for Recommendation |
|---|---|---|
Chunk size (Chunk Length) | 500–800 characters (characters) | Balances completeness of adverse event descriptions with retrieval efficiency, avoiding context loss from chunks that are too short and irrelevant information from chunks that are too long. |
Chunk Overlap Length (Chunk Overlap Length) | 50–100 characters (characters) | Ensures contextual continuity between chunks, especially when describing the progression of adverse events. |
Parsing Mode | Smart Chunking | Adapts to the mixed structured and unstructured content in pharmacovigilance documents, improving parsing accuracy. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds (seconds) | Accommodates the parsing time for large safety reports (e.g., PSURs), preventing timeouts due to oversized files. |
UPLOAD_FILE_MAX_SIZE | 500 MB | Allows uploading large R&D documents that include multiple charts and detailed appendices. |
PDF_PARSE_ENGINE | doc2x | Most pharmacovigilance documents are distributed in PDF format; this engine offers better table and text extraction performance for PDFs. |
Three Common Mistakes
- System message "Parsing failed, unable to identify file type" after file upload: This usually indicates an unsupported file format was uploaded, such as certain custom XML reports that the default system parser cannot handle.
- Key fields (e.g., drug batch number, MedDRA codes) are missing or incomplete in the parsed text: This may be due to an overly aggressive chunking strategy that breaks critical terms from their context, or the parser was not optimized for data within tables.
- API call to upload a file returns a
400 Bad Requesterror, indicating missing required parameters: This typically means thefileormetadataparameters were not passed correctly, or the request body format does not conform to the API specification.
How to Confirm Correct Configuration
- Upload and parse typical documents of different types (ADR, PSUR, RMP). Check if the parsed text content is complete, without obvious truncation, and verify the accuracy of key information extraction.
- For documents containing complex tables, verify that table content is correctly identified and converted into readable text, paying particular attention to the completeness of numerical fields like dosage and frequency, along with their units.
- Upload test files via the API interface. Check if the API response status code is
200 OKand verify that the parsing task status eventually changes to "successful." - Randomly select several parsed text chunks and examine their semantic integrity. Ensure each chunk contains independent and meaningful information, especially focusing on adverse event descriptions and causality assessment sections.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.