Data Characteristics for DTP Pharmacy Products
Core data for DTP pharmacy products originates from official drug manufacturer product inserts, internal pharmacy promotional documents, patient medication guides, and relevant regulatory files. These documents typically exist as PDFs, Word files, or scanned images. Update frequency varies from several times a month to quarterly, triggered by new drug approvals, product insert revisions, or promotional policy adjustments. Document structure for product inserts is standardized with fixed sections like "Drug Name," "Indications," and "Dosage and Administration." Promotional documents have flexible structures, often including product lists, discount information, and validity periods. Fields and units commonly include "mg" and "g" for drug dosage, "times/day" and "tablets/time" for usage, and "yuan" and "%" for prices and discounts.
Constraints from These Characteristics on Document Parsing and Chunking
The fixed section structure of drug product inserts requires parsing to identify and preserve logical segments, ensuring related information remains grouped. For example, "Adverse Reactions" and "Precautions" must remain independent chunks to prevent overlap and maintain consultation accuracy. The flexible structure and high update frequency of promotional documents demand that the parser adapts to various layouts and efficiently updates the knowledge base to reflect the latest offers. Image-based documents, especially scanned copies, require robust OCR capabilities to ensure accurate text recognition. Accurate identification and contextual association of medical terminology and measurement units are crucial for building a precise Q&A system. This avoids information loss or misinterpretation due to improper chunking.
Configuration Strategy
| Configuration Item | Recommended Value | Rationale for Recommendation |
|---|---|---|
UPLOAD_FILE_MAX_SIZE | 50 MB | Accommodates the generally large size of DTP pharmacy product inserts and related documents. |
Chunk size | 800–1200 characters | Balances information completeness and retrieval efficiency, preventing chunks from being too sparse or too dense. |
Chunk Overlap Length | 100–200 characters | Ensures contextual continuity and reduces information loss due to chunk truncation. |
PARSE_FILE_TIMEOUT_SECONDS | 300 seconds | Accounts for the time required for complex PDF parsing and OCR, preventing parsing failures due to timeouts. |
ENABLE_OCR | True | Addresses the prevalence of scanned drug inserts and image-based promotional documents. |
SPLIT_BY_HEADING | True | Leverages the fixed section structure of drug inserts to improve chunking logic. |
Three Common Pitfalls
- After uploading a file, the knowledge base remains empty, and backend logs show
Error: OCR failed. This typically indicates text recognition failure due to low image quality or improper OCR engine configuration. - When users inquire about drug dosage and administration, the response contains incorrect or missing dosage units. This usually happens when key numbers and units are separated during document chunking, leading to poor contextual relevance.
- The system cannot accurately answer questions about the latest promotional activities, even if the file has been uploaded. This often occurs because the parser fails to correctly identify the dynamic structure of promotional documents, preventing effective extraction or chunking of critical information.
How to Verify Configuration
- Select 5 random documents from different sources and formats (PDF, Word, scanned images). Upload them and check if corresponding chunks are generated in the knowledge base. Verify if the chunk content matches the original text.
- For uploaded drug inserts, check if key sections like "Indications" and "Adverse Reactions" are segmented completely and independently. Ensure no cross-section chunking or missing critical information occurs.
- Upload a promotional document containing complex tables or images. Check if table data and image text are accurately recognized and included in the chunks. Verify if retrieval performance meets expectations.
- Upload a revised drug insert via the API. Check if the corresponding old version content in the knowledge base is correctly updated or replaced. Verify the accuracy of the updated chunk content.
Note: The values provided are common starting points. Measure them against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.