Data Characteristics
DTP pharmacies are critical to the biopharmaceutical commercialization chain. Their quality documents focus on drug procurement, storage, distribution, cold chain management, and patient service processes. Data sources include regulatory documents from drug administration departments, product manuals and cold chain transportation guidelines from pharmaceutical companies, internal SOPs (Standard Operating Procedures), GSP (Good Supply Practice for Pharmaceutical Products) compliance inspection records, adverse reaction reports, and pharmacist training materials. These documents update frequently, especially with regulatory changes or new drug launches. Document structures vary, including structured tabular data (e.g., batch numbers, expiry dates, storage conditions) and extensive unstructured text (e.g., operating steps, risk descriptions, case analyses). Common and critical fields and units include drug name, batch number, production date, expiry date, temperature (℃), humidity (%RH), and storage conditions (e.g., "cool, dark place").
Constraints Imposed by Data Characteristics on Knowledge Base Retrieval and Recall
The high update frequency of DTP pharmacy quality documents requires the knowledge base to have efficient document synchronization and index update mechanisms to ensure the timeliness of retrieval results. The complex and diverse document structures, encompassing both standardized data tables and extensive free text, challenge the knowledge base's text parsing capabilities. It must accurately identify and extract key information. Diverse fields and units, especially numerical data like temperature and humidity, require the knowledge base to support numerical range queries or specific unit matching during retrieval, avoiding inaccurate recalls based solely on text similarity. Furthermore, the rigor of quality documents demands accuracy and traceability in retrieval results. Recalled content must directly link to the original source to support compliance checks and problem tracing. For retrieving specific drugs or operational procedures, the knowledge base needs to understand domain-specific terminology and the relationships between concepts for deeper semantic matching.
Configuration Settings
| Configuration Item | Recommended Value | Rationale for Recommendation |
|---|---|---|
Chunk size | 500 characters | Balances contextual information in long texts with the granularity of short texts, suiting SOPs and regulatory clauses. |
Recall count | Top 10 entries | Ensures coverage of potentially relevant documents while avoiding excessive noise, facilitating manual review. |
Similarity threshold | 0.75 | Increases similarity requirements for the rigor of quality documents, reducing inaccurate recalls. |
Rerank result count | Top 5 entries | Further optimizes relevance through a reranking model based on initial recall, improving the quality of final display. |
maxContext | 3000 token | Ensures large SOP sections or regulatory clauses are fully included in the context, preventing critical information truncation. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Allows sufficient time for parsing large PDF or Word format quality documents. |
Common Pitfalls
- Retrieval results are empty JSON or return too few items: This may be due to a
Similarity thresholdset too high, or document segments are too fragmented, preventing individual segments from fully expressing the query intent. - Recalled documents have weak relevance to the query: This may be due to
Recall countbeing set too low, failing to cover all potentially relevant documents, or a lack of effective reranking mechanisms. - After uploading documents via API, new data is not retrievable from the knowledge base: This may be due to the index update process not being triggered after document upload, or document format parsing failure, preventing data from being correctly ingested.
How to Confirm Correct Configuration
- Simulate queries against core quality management processes (e.g., cold chain management, drug acceptance) to check if recalled documents include corresponding SOPs, GSP clauses, and related records.
- Upload a document containing key numerical information (e.g., specific temperature ranges, expiry dates). Precisely query these values to verify the knowledge base's ability to accurately recall them.
- Upload a new regulatory document via the API, observe the knowledge base's index update status, and perform relevant queries immediately after the update to confirm the new document's content is retrievable.
- Compare recall results under different
Similarity thresholdsettings to determine a threshold that balances relevance with sufficient information coverage.
The values provided are common starting points and should be measured against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.