Data Characteristics in this Category
Chemistry, Manufacturing, and Control (CMC) research documents in pharmacovigilance typically contain core information about drug manufacturing processes, quality control, stability testing, and batch release. Data sources include internal research and development reports, production records, quality standard documents, batch production records, change control documents, and regulatory submissions. Document update frequency is relatively low, usually occurring during drug development or with significant process changes. Document structures are primarily mixed, including structured tables, semi-structured text, and charts. Examples include batch information tables, test result reports, and equipment calibration records. Fields include batch number, production date, expiration date, test items, analysis methods, results, units (e.g., ppm, mg/mL, %), limits, and deviation descriptions. The data volume is large and highly specialized.
Constraints from these Characteristics on Document Parsing and Chunking
The structured and semi-structured nature of CMC research documents limits traditional text segmentation methods, requiring more refined parsing strategies. Mixed charts and text require document parsers to have image and table structure recognition capabilities; otherwise, critical data may be lost. Low update frequency means initial parsing configurations must operate stably long-term. However, if document formats are subtly adjusted, rapid response and adaptation are necessary. The specialized nature of fields and units requires that context be preserved after chunking to avoid ambiguous data meanings or lost units due to truncation. For example, batch numbers, test results, and their corresponding units and limits typically form a complete semantic unit and should not be arbitrarily split. For documents like production records, time series and batch correlation are important constraints, and chunking must consider how to maintain these associations.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk Length | 500–800 characters | Balances semantic completeness and retrieval efficiency, avoiding excessive length that leads to redundancy or insufficient length that causes context loss. |
Overlap Length | 50–100 characters | Ensures semantic continuity at chunk boundaries, especially in tables or long sentence connections. |
Enable Table Recognition | Enabled | CMC documents contain extensive tabular data, such as batch information and test results, which must be recognized. |
Enable Image Text Recognition | Enabled | Some critical information may exist as images, such as chemical structures or charts, and text needs to be extracted. |
File Parsing Timeout | 600 seconds | Parsing large PDFs or documents with complex tables and images can take a long time. |
Max File Size | 200 MB | CMC reports often contain large amounts of data and charts, resulting in large file sizes. |
Three Common Mistakes
- Some tabular data in uploaded documents are not parsed, or table content is misidentified as plain text. This usually occurs because the document parser has insufficient recognition capabilities for complex table structures or tables within images, or because table recognition is not enabled.
- In the parsed knowledge base, critical batch numbers, test results, and corresponding units or limits are broken, preventing complete information retrieval during queries. This often happens when
Chunk Lengthis set too short, failing to keep complete semantic units within the same chunk. - After uploading HTML-formatted Javadoc or other API documentation, content is empty or incompletely parsed. This is because HTML file structures are complex, and the parser may fail to correctly extract valid text content or treat it as plain text, ignoring tag semantics.
How to Confirm Correct Configuration
- Select 5–10 representative CMC documents at random. Upload them and inspect the parsed chunk content to ensure critical tabular data, text in charts, and specialized terms and units are complete and accurate.
- For documents containing complex tables, verify that table rows and columns in the parsed results correspond correctly and that data matches the original document.
- Conduct simulated queries using natural language questions that include keywords like batch numbers, test items, and result ranges. Check if the retrieved chunks are accurate, relevant, and semantically complete.
- Examine document chunks for different batches and production stages within the knowledge base. Ensure they are correctly associated and differentiated, preventing data from different batches from being confused.
Note: The values given are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.