Data Characteristics in this Category
CDMO (Contract Development and Manufacturing Organization) pharmacovigilance data originates from clinical trials and post-market drug production services provided to pharmaceutical clients. Data sources include Adverse Event (AE) reports, Serious Adverse Event (SAE) reports, Periodic Safety Update Reports (PSUR), Risk Management Plans (RMP), and batch production records. Data update frequency varies based on clinical trial progress, drug market status, and regulatory requirements, potentially daily, weekly, or quarterly. Document structures are diverse, encompassing both structured tabular data and extensive unstructured text, such as handwritten doctor's notes and patient interview records. Fields include standard medical terminology, CDMO-specific batch numbers, and production line identifiers. Units involve dosage (mg, g), frequency (times/day), and time (days, weeks).
Constraints on Document Parsing and Chunking
The diversity of CDMO pharmacovigilance data poses challenges for document parsing. Unstructured text requires robust natural language processing capabilities to accurately extract adverse event descriptions, patient characteristics, and medication information. Multi-source heterogeneous data necessitates processing various file formats like PDF, DOCX, and XLSX, and unifying information structures. High-frequency data streams demand automated and incremental parsing to avoid redundant processing and resource waste. Unique fields require flexible field identification and mapping mechanisms to prevent loss of critical production and batch information. The sensitive nature of pharmacovigilance data mandates data integrity and accuracy during parsing, as any parsing error could lead to severe compliance risks. There is also a growing need to parse image content, such as identifying handwritten signatures or chart information from scanned documents.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
UPLOAD_FILE_MAX_SIZE | 500 MB | Accommodates large report files containing numerous charts or scanned documents, ensuring large file uploads are successful. |
Chunk size (Chunk Length) | 800–1200 characters (characters) | Balances contextual completeness with retrieval efficiency, suitable for longer narrative texts in adverse event reports. |
Chunk Overlap Length (Chunk Overlap Length) | 100 characters (characters) | Ensures contextual continuity at chunk boundaries, preventing critical information from being split and losing semantic meaning. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds (seconds) | Accounts for parsing time of large PDFs or documents with complex tables, preventing parsing failures due to timeouts. |
chunk_strategy | recursive_text_splitter | For unstructured text and nested structures, recursive splitting better maintains semantic integrity. |
Image OCR | Enable | Addresses text information in scanned documents and embedded images, such as handwritten notes or chart labels. |
Common Pitfalls
- Key fields (e.g.,
Drug Name,Adverse Event Description) are empty in parsing results: This typically occurs when the document parser fails to correctly identify or extract non-standard formatted fields, or when text within images is not processed by OCR. - Image content is not displayed in the preview or parsing results after uploading a PDF file: This happens when images within the PDF are not recognized as extractable objects by the parsing engine, or when OCR functionality is not enabled or improperly configured.
- Knowledge base answer accuracy is lower than expected, especially when processing DOCX or XLSX files: This may be due to a chunking strategy that is unsuitable for the internal structure of these documents, leading to critical information being inappropriately split, or tabular data not being effectively parsed into retrievable text.
How to Verify Configuration
- Select typical pharmacovigilance reports with varying formats (PDF, DOCX, XLSX) and content complexity (including images, tables, long texts). Upload them and check if the parsed chunks are complete and semantically coherent.
- Perform simulated queries against the parsed knowledge base, focusing on key information in the reports (e.g., frequency of specific adverse events, involved drug batches), and evaluate the accuracy and completeness of the answers.
- Review parsing logs to ensure there are no error messages such as
parsing timeoutorunsupported file format. Pay attention to metrics likeOCR recognition success rate. - Compare key data points from the original documents (e.g., values of the
AE_TERMfield,patient age) with the extracted information in the parsed knowledge base to ensure consistency.
The values provided are common starting points and should be measured against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.