Data Characteristics in This Category
Pharmacovigilance (PV) clinical trial pre-screening data primarily originates from clinical trial protocols, informed consent forms (ICFs), adverse event (AE) reports, serious adverse event (SAE) reports, case report forms (CRFs), and investigator brochures (IBs). These documents are largely unstructured text, predominantly in PDF and DOCX formats. AE reports require immediate submission within a defined timeframe after an event. Trial protocols and ICFs are finalized before trial initiation but may be amended during the trial, resulting in a lower update frequency. AE reports typically follow international standards such as ICH E2B, containing structured fields (e.g., event name, date of occurrence, drug name, dosage) and unstructured descriptions. CRFs contain extensive numerical and textual data, along with complex table layouts.
Constraints on Document Parsing and Chunking from These Characteristics
The immediacy and compliance requirements of pharmacovigilance data demand high efficiency and accuracy in document parsing. Key information in AE reports, such as event type and drug causality, must be precisely identified and extracted to support rapid risk assessment. The mix of structured (e.g., tabular data) and unstructured text (e.g., clinical descriptions) within documents requires parsing tools to effectively handle multiple data types while maintaining contextual integrity. For instance, tabular data in CRFs needs to preserve row and column logical relationships to prevent semantic loss from data fragmentation. Furthermore, the format diversity of documents from different sources, such as scanned PDFs or complex layouts, increases document parsing challenges. Robust Optical Character Recognition (OCR) capabilities and layout analysis techniques are necessary to ensure all information is accurately digitized.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Chunk Size) | 500-800 characters | Balances the completeness of AE descriptions with retrieval relevance, avoiding overlong chunks that dilute key information. |
Chunk Overlap Length (Overlap Size) | 50-100 characters | Maintains contextual continuity, ensuring critical information across chunks is not fragmented, supporting more precise semantic understanding. |
File Type Whitelist | pdf, docx, txt | Main formats for pharmacovigilance documents, ensuring only supported and commonly used file types are processed. |
EnabledOCR (Enable OCR) | true | Many AE reports and historical documents may be scanned, ensuring image text can be recognized and parsed. |
Table Parsing Strategy | Retain table structure and extract key rows/columns | CRFs and similar documents contain important tabular data; their structure must be maintained to prevent key numerical values from separating from descriptions. |
Max File Size | 100 MB | Accommodates clinical trial protocols and investigator brochures that contain numerous charts or scanned pages. |
Three Common Mistakes
- Incomplete table display or corrupted formatting in the knowledge base occurs when the table parsing strategy is not enabled or incorrectly configured. This leads to table content being treated as plain text and incorrectly chunked.
- Retrieval results lack detailed descriptions of critical adverse events. This happens when relevant information is truncated or scattered across multiple discontinuous chunks, due to a
Chunk size(Chunk Size) that is too short and insufficientChunk Overlap Length(Overlap Size), leading to loss of important context. - Some content cannot be retrieved after document upload. This typically happens when the document is a scanned image, but the
EnabledOCR(Enable OCR) function is not activated, preventing text in the image from being recognized.
How to Verify Correct Configuration
- Upload typical adverse event reports and case report forms. Check if the chunked content in the knowledge base fully retains critical drug names, event descriptions, and dosage information.
- Use the knowledge base preview function to review documents containing complex tables. Ensure that the logical relationships of row and column data in tables are correctly maintained, and that key numerical values are not separated from their corresponding descriptions.
- Upload a PDF document known to contain scanned text. After retrieval, verify that the scanned text content can be correctly recalled to confirm that the
EnabledOCR(Enable OCR) function is working effectively.
Note: The values provided are common starting points. Measure against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.