Data Characteristics in this Domain
Pharmacovigilance regulatory submission documents primarily include Individual Case Safety Reports (ICSRs), Periodic Safety Update Reports (PSURs/PBRERs), and Risk Management Plans (RMPs). These documents are typically in PDF format, feature complex internal structures, and contain extensive unstructured text, semi-structured tables, and charts. Data sources often include global clinical trials, post-market surveillance, literature reviews, and case reports. Regarding update frequency, ICSRs are real-time, PSURs/PBRERs are usually submitted semi-annually or annually, and RMPs are updated as needed. Document fields include patient demographics, drug information, adverse event descriptions, medical terminology (e.g., MedDRA codes), dosage, administration, event occurrence time, and outcomes. Units involve time (days, weeks, months) and dosage (mg, g, ml), and cross-language descriptions are common.
Constraints Imposed by these Characteristics on Document Parsing and Chunking
The complex document structure and diverse sources of pharmacovigilance data place specific demands on document parsing and chunking. First, nested tables and charts within PDFs require specialized handling to ensure data integrity and prevent parsing into meaningless text blocks. Second, accurate identification and contextual association of medical terminology and specialized abbreviations (e.g., MedDRA codes) are crucial for understanding the nature and severity of adverse events. Mixed-language text, especially in adverse event descriptions, requires the parser to have multilingual recognition capabilities. Furthermore, the real-time nature of ICSRs means the parsing system needs to support rapid processing and incremental updates, while the periodic updates of PSURs/PBRERs require version management capabilities to ensure accurate comparison of information between different versions. Precise identification of fields and units, particularly dosage and time, directly impacts the accuracy of safety assessments. Therefore, chunking must maintain a close connection between relevant values and their descriptions.
Configuration Guidelines
| Configuration Item | Suggested Value | Rationale |
|---|---|---|
Chunk size | 500–800 characters | Balances semantic completeness with recall efficiency. Avoids overly long texts diluting key information or overly short texts losing context. |
Chunk Overlap Length | 100–150 characters | Ensures semantic continuity at chunk boundaries, especially when tables or long sentences span across chunks. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Prevents timeout interruptions when processing large documents like PSURs/RMPs, which can take extensive parsing time. |
table_parsing_strategy | auto or advanced | Automatically identifies and structurally parses complex tables, preserving row and column relationships, and improving table data usability. |
ocr_languages | zh, en | Addresses potential mixed Chinese and English content in documents, ensuring text in images or scanned documents is recognized. |
recall_threshold | 0.75 | Improves recall precision for medical terminology and adverse event descriptions, reducing interference from irrelevant information. |
Three Common Mistakes
- Parsing results contain numerous fragmented table rows or columns, making data unintelligible. This occurs because
table_parsing_strategyis not configured correctly, or the document's table structure is too complex for effective boundary recognition. - Key medical terms or dosage units are separated from their descriptive information after chunking, leading to a lack of context during retrieval. This happens if
Chunk sizeis too small, cutting closely related information, or ifChunk Overlap Lengthis insufficient. - Uploading large PDF documents results in
Request Timeoutor504 Gateway Timeouterrors. This happens if thePARSE_FILE_TIMEOUT_SECONDSparameter is set too low, not allowing enough processing time.
How to Verify Configuration
- Randomly select multiple types of pharmacovigilance documents. Upload them and review the chunked content in the knowledge base. Check if key information (e.g., MedDRA codes, dosage, adverse event descriptions) is complete and contextually coherent.
- For PDF files containing complex tables, verify that the parsed chunks clearly display the table content, or that table data is structurally extracted. This can be validated by retrieving specific data from the tables.
- Monitor log output to ensure no significant timeout or parsing failure errors occur during processing. Verify the actual effect of
PARSE_FILE_TIMEOUT_SECONDS. - Attempt to retrieve adverse event descriptions from the knowledge base. Evaluate the accuracy and relevance of the recall results, and adjust
recall_thresholdbased on actual business needs.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.