Data Characteristics in Rational Drug Use
Data for rational drug use registration documents primarily originates from drug inserts, clinical trial reports, pharmacopoeia standards, drug interaction databases, adverse event reports, domestic and international guidelines, and relevant regulatory files. These documents typically exist as PDFs, Word files, or structured databases (e.g., XML). Data updates frequently, especially with drug insert revisions, new drug approvals, and guideline updates. Document structures are complex, containing extensive specialized terminology, dosage units (e.g., mg, g, IU), administration routes, indications, contraindications, drug metabolism pathways, and other information. Document formats and layouts vary significantly across sources.
Constraints Imposed by Data Characteristics on Vector Models and Indexing
The specialized and complex structure of rational drug use data demands high semantic understanding from vector models. Extensive professional vocabulary and medical abbreviations require models sensitive to domain knowledge to avoid losing critical information during chunking and vectorization. High update frequency necessitates an indexing system that supports efficient incremental updates and version management, ensuring the timeliness of retrieval results. Diverse document formats and mixed structured data make text extraction and structured information recognition critical challenges during preprocessing, impacting the accuracy of the final vector representation. Numerical information, such as dosages and units, requires special handling during vectorization to preserve its quantitative meaning.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale for Recommendation |
|---|---|---|
Chunk size (Chunk Length) | 800–1200 characters | Balances contextual completeness with vector model processing efficiency, ensuring each chunk contains sufficient semantic information. |
Overlap Length | 100–200 characters | Ensures semantic continuity at chunk boundaries, preventing critical information from being cut off. |
Recall count (Recall Count) | Top 8–15 items | Covers a broader range of potentially relevant document snippets, improving retrieval recall. Re-ranking refines results later. |
Similarity threshold (Similarity Threshold) | Calibrate by empirical measurement | Adjusts based on actual query performance, false positive rate, and false negative rate. Typically between 0.7–0.85. |
embedding_model | bce-embedding-v1 or text-embedding-ada-002 | Prioritize models performing well in the medical domain or those with sufficiently large context windows. |
maxContext | 32000 tokens | Ensures enough recalled content fits when generating responses, preventing information loss due to context truncation. |
Common Pitfalls
- After uploading many documents, the index status remains "Indexing" for an extended period, but actual progress is slow or stalled. This often results from the file parser encountering errors when processing complex or corrupted PDF files, leading to a blocked task queue, or
PARSE_FILE_TIMEOUT_SECONDSbeing set too short. - Retrieval results lack highly relevant drug dosage or specific indication information for a query. This may relate to
Chunk size(Chunk Length) being set too large, causing a chunk to contain too much irrelevant information that dilutes key semantics, orOverlap Lengthbeing too short, leading to loss of critical information at chunking points. - An
error: { message: 'This Token Is Not Authorized To Use The Model:text-embedding-3-large (request id: 20240...' }error occurs when attempting to callembedding_model. This indicates the configured API Key lacks permission to access the specified Embedding model. Check the API Key's authorization scope or use an available key.
Verification of Configuration
- Upload different types of rational drug use documents (PDF, DOCX, XML). Verify that indexing tasks complete normally and check the index status.
- Execute retrievals for core queries such as drug indications, contraindications, and interactions. Check if the returned results include accurate and complete key information, like drug names, dosage units, and mechanisms of action.
- Use queries containing specialized medical terminology and abbreviations. Observe the quality and relevance of recall results to ensure the vector model accurately understands domain-specific vocabulary.
- Adjust
Similarity threshold(Similarity Threshold) andRecall count(Recall Count). Observe changes in the quantity and relevance of retrieval results to find a balance that reduces interference from irrelevant information while maintaining high recall.
The values provided are common starting points and should be measured against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.