Data Characteristics
Data in the medical affairs pharmacovigilance domain originates from clinical study reports, post-market surveillance data, regulatory guidelines and announcements, and individual case safety reports (ICSRs). These documents update frequently, especially when new drugs launch or new adverse event signals emerge. Document structures are complex, often containing a mix of structured and unstructured data like medical terminology, dosage information, patient characteristics, adverse event descriptions, and causality assessments. Fields and units are highly specialized; for instance, "dosage" might be in milligrams (mg), micrograms (mcg), or millimoles (mmol), and "adverse event incidence" often appears as a percentage or per thousand person-years (PY). Many reports are in PDF format, potentially including tables, charts, and scanned images, which challenges text extraction.
Constraints Imposed by These Characteristics on Document Parsing and Chunking
High update frequency demands efficient and automated document parsing to quickly ingest and process newly published medical literature. Complex document structures, particularly tables and scanned images, mean that simple text extraction is insufficient. Deep parsing requires layout analysis and OCR technology. Specialized medical terminology and diverse unit systems necessitate chunking that preserves term integrity, preventing truncation of critical information. For example, the adverse event description severe liver dysfunction with jaundice should not be split. Documents often contain numerous citations and footnotes; these require special handling during chunking, typically separating them from the main discussion to avoid interfering with the recall quality of core information. Identifying and independently chunking key sections like causality assessments is crucial for subsequent knowledge retrieval and reasoning.
Configuration Guidelines
| Configuration Item | Suggested Value | Rationale for This Value |
|---|---|---|
chunk_size | 800–1200 characters | Balances the paragraph length of medical reports with contextual completeness, preventing critical information from being split. |
overlap_size | 100–200 characters | Ensures sufficient overlap between adjacent chunks to capture cross-paragraph related information, especially for causality chains. |
maxContext | 8192 | Accommodates the long-text nature of medical documents, providing a sufficient context window for accurate medical terminology understanding and reasoning. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Accounts for the parsing time of large clinical study reports and complex PDFs, preventing parsing failures due to timeouts. |
ocr_enabled | true | Many medical documents contain scanned images or embedded images; enabling OCR ensures all text content can be extracted. |
table_parsing_strategy | auto | Automatically identifies and structurally extracts table data, used for processing structured information like dosage and patient characteristics. |
Three Common Mistakes
- Important medical terms are truncated or split in parsing results. This occurs because
chunk_sizeis too small, failing to preserve complete concepts within a chunk. - After importing PDF documents, some table data is not correctly recognized or is empty. This happens when
ocr_enabledis not active ortable_parsing_strategyis misconfigured, failing to handle complex table layouts. - When processing large clinical trial reports, parsing frequently times out. This is due to
PARSE_FILE_TIMEOUT_SECONDSbeing set too short, not allowing enough time for file parsing to complete.
How to Verify Configuration
- Select a typical medical report containing complex tables, scanned text, and multiple pages. Upload it and examine the parsed knowledge base chunks to ensure all critical information is accurately extracted and not truncated.
- Perform retrieval tests on the parsed knowledge base. Use core medical terms and adverse event descriptions from the report as queries to verify that the recalled chunks contain complete contextual information.
- Check system logs to confirm no
timeoutorparsing errormessages appear when processing various medical documents. - Compare key fields (e.g., drug name, dosage unit, adverse event name) before and after parsing to ensure the accuracy and completeness of the parsed data meet expectations.
The values provided are common starting points and should be measured against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.