Data Characteristics
Data for DTP pharmacy product and reagent consultations primarily originates from drug inserts, pharmaceutical company product manuals, clinical research reports, patient education materials, and common Q&A compiled by pharmacists. This data has a relatively stable update frequency, with major updates occurring quarterly or semi-annually when new drugs are launched or drug inserts are revised.
Document structures are typically well-defined. Drug inserts and product manuals often have standardized sections, including fixed fields like ingredients, indications, dosage and administration, and adverse reactions. Clinical reports are even more structured, containing research objectives, methods, results, and conclusions. Key fields include drug name, generic name, dosage form, specifications, manufacturer, and approval number, often with time-sensitive information such as batch number and expiration date.
Constraints Imposed by Data Characteristics on Vector Models and Indexing
The structured nature of DTP pharmacy data allows for sophisticated pre-processing before vectorization. Identifying key entities like drug names and indications can significantly improve the precision of vector recall. The low update frequency means that the cost of rebuilding or incrementally updating the vector library is manageable, eliminating the need for frequent full re-indexing.
However, documents contain numerous specialized terms and medical abbreviations, demanding a high level of semantic understanding from the vector model. Time-sensitive information, such as drug batch numbers and expiration dates, is not directly involved in vectorization but requires additional matching and filtering during result display to prevent providing outdated or invalid information. Furthermore, user queries may include colloquial descriptions or symptoms, requiring the vector model to bridge the semantic gap between professional terminology and everyday language.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Segment Length) | 500–800 characters (characters) | Balances semantic completeness with vector model processing efficiency, avoiding information overload or fragmentation in a single segment. |
Chunk Overlap Length (Segment Overlap Length) | 50–100 characters (characters) | Ensures contextual continuity and minimizes information loss at segment boundaries. |
Recall count (Recall Count) | 8–12 entries (items) | Covers sufficient relevant information while controlling computational load for subsequent re-ranking and generation stages. |
Similarity threshold (Similarity Threshold) | Calibrate based on actual measurements | Balances recall and precision according to business needs; start adjusting from around 0.75. |
Rerank result count (Re-ranked Return Count) | 3–5 entries (items) | Further filters the most relevant document snippets, reducing the burden on the generation model and improving response speed. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds (seconds) | Accommodates parsing time for large drug inserts or clinical reports, preventing timeouts. |
Common Pitfalls
- After creating a new collection, the indexing status remains stuck at "processing" or "indexing last batch" for an extended period. This often results from document parsing timeouts or insufficient memory, especially when processing large PDF or image-based inserts.
- Vector retrieval results deviate significantly from expectations, sometimes showing irrelevant content. This can occur if the vector model fails to adequately understand specialized medical terminology, leading to inaccurate semantic matching.
- The initial retrieval response time is excessively long, exceeding
8 seconds(seconds), for example. This might be due to insufficient optimization of the vector database index or hardware resource limitations (e.g., SSD performance) affecting retrieval speed.
How to Verify Configuration
- Upload typical drug inserts and product manuals. Check if the
indexing statusshows "completed" and confirm that indexing time is within a reasonable range. - Conduct multiple rounds of
semantic retrievaltests for common drug consultation questions. Observe if the returnedRecall count(recall count) andsimilarity scoreare reasonable, and manually evaluate the relevance of the top few recalled results. - Simulate user queries, including both professional terms and colloquial expressions. Check if the information in the
Rerank result count(re-ranked return count) is accurate and answers the questions. Compare results under differentSimilarity threshold(similarity thresholds). - Monitor system resource usage, such as CPU, memory, and disk I/O. Ensure that
vector retrievalresponse times remain stable under peak consultation pressure, avoiding delays exceeding10 seconds(seconds).
The values provided are common starting points and should be measured against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.