Data Characteristics
DTP pharmacies prepare registration documents using data from several sources: original submission materials from pharmaceutical manufacturers, regulations and guidelines from national medical product administrations, technical review reports, and the pharmacy's own operational qualifications and quality management system documents. Data updates are infrequent, typically occurring when regulations change or new drugs are submitted for approval. Document structures are complex. They include many technical PDFs, WORD approval documents, EXCEL reports, and scanned images. Fields and units are highly specialized. Examples include drug generic names, brand names, dosages, specifications, batch numbers, expiration dates, registration numbers, and approval numbers. Units use medical measurements such as mg/ml, IU, and g/L, often with abbreviations and aliases.
Constraints on Knowledge Base Retrieval and Recall
The complex data characteristics of DTP pharmacy registration documents impose specific constraints on knowledge base retrieval and recall. Diverse document formats require robust file parsing, especially for accurate table and image text recognition within PDFs. Infrequent updates necessitate an efficient version management mechanism to ensure retrieval of the latest valid regulations or approvals. Recognizing specialized fields and units is a core challenge. Traditional keyword matching often misses synonyms or abbreviations, requiring semantic understanding and entity recognition technologies. Documents contain many specific codes like batch numbers and registration numbers. Retrieval must support precise queries based on structured metadata, not just content, to avoid confusing different drug documents.
Configuration Settings
| Configuration Item | Suggested Value | Rationale |
|---|---|---|
Chunk size | 800–1200 characters | Paragraphs in DTP pharmacy documents often have long semantic integrity. Shorter segments can lose context, while longer segments add irrelevant information. |
Recall count | Top 10 entries | Submission documents are highly interconnected. Increasing recall items covers more potentially relevant content and improves hit rates. |
Similarity threshold | Calibrate by actual measurement | This avoids missing recalls due to variations in professional terminology and prevents recalling too many irrelevant results. |
Rerank result count | Top 5 entries | After initial filtering by recall, reranking further prioritizes the most relevant information. |
UPLOAD_FILE_MAX_SIZE | 500 MB | Submission documents frequently contain large PDF files, this ensures unrestricted uploads. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Parsing complex scanned documents or multi-layered nested PDFs can be time-consuming. |
Common Misconfigurations
- "File size exceeds limit" when uploading to the knowledge base: The
UPLOAD_FILE_MAX_SIZEparameter is set too low for large submission documents. - Retrieval results contain much irrelevant information or miss key details: Similarity scores are low. This indicates an unreasonable segmentation strategy or an unoptimized
Similarity threshold. - The knowledge base interface freezes or is unresponsive for a long time after clicking: Requests time out. The
PARSE_FILE_TIMEOUT_SECONDSis set too short, preventing complex document parsing tasks from completing within the allotted time.
Verification Steps
- Upload typical submission document files. Check upload time and system resource usage to ensure they are within acceptable limits.
- Design search terms for different drugs and regulatory clauses, including synonyms, abbreviations, and specific codes. Verify the accuracy and completeness of recall results, checking if all expected relevant documents are included.
- Randomly select documents from the knowledge base. Use their content as search terms for reverse queries. Verify if the corresponding documents are effectively recalled and evaluate their ranking in the results.
- Simulate complex problems from actual submission scenarios. Conduct multi-turn dialogue tests. Observe if the knowledge base maintains high-quality retrieval and recall performance under continuous questioning, especially its ability to integrate information across documents.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.