Data Characteristics
DTP pharmacy quality documents include pharmaceutical management policies, drug maintenance records, temperature and humidity monitoring reports, supplier qualification certificates, batch inspection reports, and adverse reaction records. These documents are often in Word, PDF, or Excel formats, with some scanned images embedded. Data updates are frequent; for example, temperature and humidity records update daily, while drug batch information and supplier qualifications update with procurement and audit cycles. Document structures are relatively standardized, typically featuring fixed heading levels, section numbers, and table layouts. Fields involve drug generic names, batch numbers, expiry dates, manufacturers, storage conditions, and testing indicators with units (e.g., mg/tablets, ℃, %RH), requiring precise identification.
Constraints Imposed by These Characteristics on Document Parsing and Chunking
The structured nature of DTP pharmacy documents requires parsers to accurately identify headings, paragraphs, and tables to maintain semantic integrity. Frequent data updates, especially for temperature/humidity records and batch information, demand efficient re-parsing and index updating. Embedded images (e.g., scanned documents, charts) require OCR capabilities, ensuring text content links to the image context. Precise extraction of critical fields like drug batch numbers and expiry dates is fundamental for accurate subsequent Q&A; inaccurate parsing can lead to incorrect drug information. Additionally, different document types (policies, reports, records) should use distinct chunking strategies to avoid mixing semantically different paragraphs.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
UPLOAD_FILE_MAX_SIZE | 500 MB | A single DTP pharmacy quality document (e.g., an annual audit report) can contain numerous images and attachments, leading to large file sizes. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Large PDF or Word documents, especially those with complex tables and multi-level structures, require longer parsing times. |
Chunk size | 800–1200 characters | Ensures each chunk contains sufficient context while avoiding excessive length that could lead to information redundancy and affect recall accuracy. |
Chunk Overlap Length | 100 characters | Maintains semantic continuity between chunks, preventing critical information from being truncated. |
ENABLE_OCR | True | Quality documents often include scanned copies or image-based batch inspection reports and signature pages, requiring text recognition from images. |
TABLE_PARSING_STRATEGY | markdown | Data such as drug batches and temperature/humidity records are often presented in tables; converting to Markdown facilitates subsequent understanding. |
Common Pitfalls
- Symptom: After uploading some documents, critical drug batch information in tables cannot be recalled during conversations. Reason: The
TABLE_PARSING_STRATEGYwas not set tomarkdownortext, causing table content to be ignored or incorrectly parsed. - Symptom: After uploading an updated temperature and humidity record document, the model still provides old data or returns a
408 Request Timeouterror. Reason: File size or parsing complexity exceeded thePARSE_FILE_TIMEOUT_SECONDSsetting, causing parsing to fail, or the document re-indexing was not triggered. - Symptom: When discussing image content within a document, the model cannot provide relevant information. Reason:
ENABLE_OCRwas not enabled, or the OCR service recognition rate was insufficient, preventing text content from images from being extracted and indexed.
Verification Steps
- Upload a quality document containing complex tables and scanned images. Check its indexing status. Use the
GET /api/v1/vectors/searchinterface to query the document content and verify that tables and image text are correctly chunked and indexed. - For specific drug batch numbers or expiry dates, perform Q&A tests to verify that the model accurately recalls relevant information and that the
chunk_idin the recall results points to the correct document area. - Upload a frequently updated document (e.g., temperature and humidity records). Observe whether Q&A results reflect the latest data promptly after re-uploading, and check the
last_modifiedfield's update status.
Note: The values provided are common starting points. They should be measured against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.