Data Characteristics
Rational drug use data primarily originates from various policy documents, technical guidelines, and expert consensuses published by national health and drug administration bodies. It also includes internal regulations and Standard Operating Procedures (SOPs) from medical institutions. Data updates are relatively stable: policies and regulations typically update annually or in response to major events, while guidelines and consensuses may update every few months to a year.
Document formats vary, encompassing official PDF files, Word documents, HTML pages, and some image or table-based drug catalogs and contraindication lists. Document structures often include chapters, sections, and clauses, potentially with charts and appendices.
Fields and units involve drug names, generic names, dosages, administration routes, indications, contraindications, adverse reactions, interactions, and normal/abnormal thresholds for clinical indicators (e.g., mg/kg, IU, mol/L).
Constraints on Vector Models and Indexing
Policy and guideline documents for rational drug use are often lengthy and complex, containing specialized terminology and logical connections. Vector models must capture long-range dependencies and semantic relationships between concepts.
The moderate update frequency requires an indexing strategy that balances initial build efficiency with convenient incremental updates. Non-structural information like images and tables in documents needs additional processing (e.g., OCR or table parsing) to convert content into vectorizable text descriptions.
Diverse fields and units, especially numerical clinical indicators, may require range queries or unit conversions during retrieval. This demands the vector index to handle structured information or standardize it during preprocessing. Furthermore, cross-references and interconnections between regulatory clauses increase the need for complete contextual retrieval results.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Segment Length) | 800–1200 characters | Ensures each segment contains sufficient context while avoiding excessive length that could lead to semantic drift. |
Chunk overlap (Segment Overlap) | 100 characters | Connects adjacent segments, improving retrieval recall and reducing context breaks. |
embedding_model | text-embedding-ada-002 or bge-large-zh | Balances semantic understanding with Chinese language processing, adapting to specialized terminology. |
Recall count (Recall Count) | Top 5–8 items | Covers highly relevant original document snippets, providing enough candidates for subsequent reranking. |
Similarity threshold (Similarity Threshold) | Calibrated based on actual measurements | Balances precision and recall based on the specific dataset and model performance. |
Rerank result count (Rerank Return Count) | Top 3 items | Selects the most relevant few snippets, improving the accuracy and conciseness of the final answer. |
Common Pitfalls
- Issue: Retrieval results contain many irrelevant or low-relevance document snippets, failing to answer specific questions about drug dosages or contraindications. Reason:
Chunk size(Segment Length) is set too long, causing information within a single segment to be too dispersed and the vector representation to be unfocused. Alternatively,Similarity threshold(Similarity Threshold) is set too low, recalling too many weakly relevant results. - Issue: Image content in some documents, such as illustrations or tables in drug inserts, cannot be retrieved. Reason: During file upload and processing, image content was not effectively OCR-identified or table-structured, preventing its information from being converted to text and vectorized.
- Issue: After local FastGPT deployment, the model or vector service connection fails, reporting
Connection refusedorTimeout. Reason: The model address or API key in theoneapiconfiguration is incorrect, or a firewall is blocking network communication between the FastGPT container and the model service.
Validation Steps
- Upload typical rational drug use regulation documents. Verify that each segment in the knowledge base is complete, semantically coherent, and does not truncate important information.
- Test with questions of varying granularity (e.g., "indications of a certain drug," "medication guidelines for a certain disease"). Conduct multiple rounds of testing, compare the
similarityscores of retrieval results, and manually assess the relevance of recalled snippets to the questions. - Use documents containing images and tables for testing. Verify that their content is correctly parsed, vectorized, and effectively recalled during retrieval.
The values provided are common starting points and should be measured against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.