Data Characteristics
DTP pharmacy product and reagent consultation data originates from manufacturer-provided product inserts, drug approval documents, clinical trial reports, pharmacological and toxicological studies, and patient medication guides. These documents are typically in PDF, Word, or structured database formats. Data updates are frequent, especially with new drug approvals, expanded indications, or adverse event monitoring reports, potentially updated weekly or even daily. Document structures are rigorous, containing extensive specialized terminology, dosage units, contraindications, and interactions. Common fields include generic name, brand name, approval number, manufacturer, dosage form, specifications, indications, dosage and administration, adverse reactions, precautions, drug interactions, and storage. Units encompass various measurements such as mg, ml, IU, g/L, and times/day.
Constraints Imposed by Data Characteristics on Knowledge Base Retrieval and Recall
The specialized and rigorous nature of DTP pharmacy data places high demands on knowledge base retrieval and recall. First, high update frequency requires the knowledge base to support rapid incremental updates and version management to ensure the timeliness and accuracy of query results. Second, the presence of numerous specialized terms and measurement units in documents requires the tokenizer and retrieval model to accurately identify and understand them, preventing inaccurate recall due to tokenization errors. For example, "three times a day, one pill each time" requires "three times a day" to be processed as a single semantic unit. Furthermore, the complex structure of drug inserts, such as multi-level headings, tables, and footnotes, demands robust structured information extraction capabilities from the document parser to ensure content completeness and prevent the omission of critical information. Simultaneously, user inquiries often involve cross-referencing multi-dimensional information, such as "Can XX drug be taken with YY drug?", which necessitates the knowledge base to possess associative retrieval capabilities to extract and integrate relevant snippets from different documents.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Chunk Length) | 500–800 characters (characters) | Accommodates the paragraph length of drug inserts, balancing semantic completeness with retrieval efficiency. |
Chunk overlap (Chunk Overlap) | 50–100 characters (characters) | Ensures critical information (e.g., dosage, administration) spanning across chunks is not cut off during retrieval. |
Recall count (Recall Count) | Top 5–8 entries (top 5–8 items) | Covers multiple aspects a user might be interested in, such as indications, adverse reactions, and dosage. |
Similarity threshold (Similarity Threshold) | Calibrated by actual measurement | Adjusts based on actual query performance to balance recall rate and accuracy, avoiding interference from irrelevant information. |
Rerank result count (Rerank Return Count) | 3–5 entries (3–5 items) | Reranks initial recall results to prioritize the most relevant and authoritative information. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds (seconds) | Addresses potentially long parsing times for large PDF files like drug inserts. |
Three Common Pitfalls
- Some content is lost after importing knowledge base documents, such as table data or descriptions below images. This occurs because the document parser fails to correctly identify and extract all information from complex structures.
- After a user asks a question, the AI's answer contradicts facts in the knowledge base or exhibits "hallucinations." This likely happens when too few or low-quality relevant snippets are retrieved, failing to provide sufficient background information for the answer.
- The knowledge base training status remains "training" for an extended period, or query results do not reflect updates promptly after data changes. This could be due to a congested backend task queue or an incorrectly triggered incremental update mechanism.
How to Verify Configuration
- Select a batch of DTP pharmacy documents containing specialized terminology, dosage units, and complex structures. Import them and check if the preview content is complete, especially tables and footnote information.
- For core fields like generic drug names, indications, dosage and administration, and adverse reactions, construct multiple test questions including fuzzy and exact queries. Verify the accuracy and relevance of retrieval results and adjust the
Similarity threshold(Similarity Threshold) based on actual needs. - Simulate scenarios like new drug launches or drug insert updates. Upload the updated documents, check if the knowledge base update process is smooth, and verify that updated query results reflect the latest data.
- Check the log system to confirm no abnormal errors or prolonged timeouts occur during document parsing, vectorization, and retrieval, especially regarding the
PARSE_FILE_TIMEOUT_SECONDSparameter.
The values provided are common starting points and should be measured against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.