Data Characteristics in this Domain
GMP-compliant pharmacovigilance data primarily originates from various reports and documents submitted by Marketing Authorization Holders (MAHs). These include Adverse Drug Reaction (ADR) reports, Periodic Safety Update Reports (PSURs), Risk Management Plans (RMPs), and change control documents. These documents are typically in PDF, Word, or scanned image formats, containing extensive structured and unstructured information. Data updates occur frequently; ADR reports may be submitted in real-time or periodically, while other regulatory documents are updated regularly according to regulatory requirements. Document structures are complex, often including tables, figures, batch information, dosage units, patient characteristics, and descriptions of drug interactions. They strictly adhere to templates and terminology mandated by regulatory bodies. Fields and units are precise; for example, drug dosages are often exact to milligrams (mg) or International Units (IU), ages are precise to years, months, or days, and adverse reaction codes follow the MedDRA dictionary.
Constraints Imposed by these Characteristics on Document Parsing and Chunking
The complexity and standardization of GMP-compliant pharmacovigilance documents impose specific requirements on document parsing and chunking. First, the numerous tables and figures within documents require parsers to accurately identify table boundaries, rows, and columns, extract internal data, and understand accompanying textual descriptions to ensure no critical information is lost. Second, strictly defined fields and units in regulatory reports demand that related fields (e.g., "drug name," "dosage," "unit") are processed as a logical whole during chunking to prevent semantic fragmentation. High update frequency means the knowledge base must support incremental updates and version management. Chunking strategies should adapt to minor changes in document content, ensuring updated knowledge chunks accurately link to older data. Finally, strict terminology and coding systems (like MedDRA) require maintaining the integrity of these specialized terms during chunking, avoiding semantic disruption due to splitting, which could affect subsequent retrieval and association accuracy.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size | 500–800 characters | Balance context completeness with avoiding overly large chunks, especially for adverse reaction descriptions. |
Overlap Length | 100–150 characters | Ensure sufficient contextual overlap between adjacent chunks to handle cross-paragraph related information, such as dosage and administration route. |
EnabledTable Parsing | Yes | GMP documents contain a large amount of critical tabular data, such as batch information and adverse reaction lists, which must be fully parsed. |
ParsingTimeout | 600 seconds | Processing large PSUR or RMP files can take a long time; this prevents parsing failures due to oversized files. |
Vector Model | text-embedding-ada-002 | This model effectively understands specialized terminology and complex semantic relationships in the biomedical field. |
Text Preprocessing Rules | Custom Regex | Custom regular expressions can preprocess specific formats like MedDRA codes and drug batch numbers, improving chunking quality. |
Three Common Mistakes
- After importing tabular datasets into the knowledge base, retrieval results do not display table content, and logs show "table data not associated with text blocks." This occurs when table parsing results are not effectively linked to surrounding text or not correctly extracted as independently retrievable fields.
- Uploaded documents appear as multiple disjointed small chunks in the knowledge base, making it impossible to obtain complete context during retrieval. This happens when
Chunk sizeis set too small or when the characteristics of long paragraphs and complex sentences in regulatory documents are not adequately considered. - Documents chunked locally using vector model A are uploaded, but the server uses vector model B, leading to poor retrieval performance. This is because different vector models generate embeddings in different vector spaces; parsing and retrieval must use consistent vector models.
How to Confirm Proper Configuration
- Select typical adverse reaction reports and PSUR documents. Check if the parsed chunks contain all key information, especially tabular data and specialized terminology.
- Perform retrievals for specific drug names, adverse reaction descriptions, or batch numbers. Verify that the returned knowledge chunks are complete and semantically coherent in context, then evaluate recall rate.
- Examine randomly sampled document chunks in the knowledge base. Confirm that
Overlap Lengtheffectively connects adjacent semantics, preventing critical information from being cut off.
The values provided are common starting points and should be measured against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.