Data Characteristics
Pharmacovigilance data for small molecule drugs primarily originates from clinical trial reports, post-marketing surveillance reports (e.g., CIOMS I forms), drug labels, medical literature, and regulatory safety updates. These documents are typically in PDF format, with a smaller number of Word or structured data files. Clinical trial reports are usually released after trial completion. Post-marketing surveillance reports accumulate continuously based on adverse event occurrences. Regulatory updates are irregular.
Document structures for reports often include abstracts, main bodies, figures, tables, and appendices. Information such as adverse event descriptions, drug dosages, administration routes, and patient characteristics are distributed across different sections. Fields and units are highly standardized, including dosage (mg, g), frequency (times/day, QD), time (days, weeks), adverse event terms (MedDRA codes), and laboratory indicators (mmol/L, U/L).
Constraints on Document Parsing and Chunking
The characteristics of small molecule drug pharmacovigilance documents impose specific requirements on document parsing and chunking. Primarily, the prevalence of complex PDF formats demands parsers that accurately identify text, tables, and images, and handle multi-column layouts. Adverse event descriptions may span pages or be embedded in long paragraphs, requiring intelligent chunking strategies to maintain semantic integrity.
Frequent updates to surveillance reports emphasize the importance of incremental parsing and version management to avoid reprocessing already parsed content. The high standardization of fields and units allows for more precise extraction of key information after chunking, such as 50 mg dosage or elevated liver enzymes events, using regular expressions or named entity recognition. Additionally, the specialized nature of medical terminology requires chunking to avoid splitting key phrases, ensuring complete context during retrieval. For tabular data, the parser must convert it into queryable structured text for subsequent retrieval.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size | 800–1200 characters | Balances contextual completeness with retrieval efficiency; avoids splitting critical adverse event descriptions. |
Chunk Overlap Length | 100–200 characters | Ensures key information spanning across chunks is not lost, improving recall quality. |
PARSE_FILE_TIMEOUT_SECONDS | 300–600 seconds | Handles large clinical reports or PDFs with many figures, preventing parsing timeouts. |
maxContext | 32000 token | Accommodates lengthy medical reports, providing large language models with sufficient context for complex cases. |
PDF_OCR_ENABLED | true | Ensures content from scanned or image-based reports can be recognized and parsed. |
CHUNK_STRATEGY | By Title and Paragraph | Prioritizes maintaining the document's logical structure, such as adverse event sections and dosage information sections. |
Common Pitfalls
- Symptom: After uploading a PDF file, some table content is not extracted correctly, leading to missing key dosage or laboratory data. Cause: The parser fails to correctly identify complex table structures in the PDF, confusing their content with regular text or skipping them entirely.
- Symptom: System logs show
PDF parsing failed: doc2x error, and file parsing is interrupted. Cause: The PDF file may be encrypted, corrupted, or have an abnormal internal structure, preventing the underlying parsing tool from processing it correctly. - Symptom: The large language model's response fails to mention an adverse reaction event clearly present in the original text, or the event description is incomplete. Cause: During document chunking, critical event descriptions are improperly split, leading to individual chunks lacking sufficient contextual information.
Verification Steps
- Select a typical small molecule drug clinical trial report with complex tables and multi-column layouts. Upload it and examine the parsing results to confirm that table data is complete and structured.
- Choose several pharmacovigilance reports from different sources (e.g., regulatory agencies, pharmaceutical companies). Upload them and review the chunk preview to confirm that key information like adverse events and dosages maintain semantic integrity within reasonable chunks.
- For PDF files that fail to parse, try enabling the
PDF_OCR_ENABLEDparameter and re-uploading to observe if parsing results improve.
The values provided are common starting points and should be measured against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.