Vector Model and Index for Off-Label Drug Use Medical Information (MI) Response

Off-label drug use medical information primarily originates from clinical research literature, case reports, expert consensus, international

Data Characteristics

Off-label drug use medical information primarily originates from clinical research literature, case reports, expert consensus, international guidelines, and extended interpretations of drug labels. Data update frequencies vary: clinical research and guidelines may update quarterly or annually, while case reports might publish irregularly. Document structures are diverse, including structured database records, semi-structured clinical trial reports (PDF format), and unstructured text descriptions. Common fields include drug name, indication, dosage and administration, adverse reactions, mechanism of action, and level of supporting evidence. Units involve dosage (mg, g, IU), time (hours, days, weeks), and frequency (times/day), with potential for mixed units or abbreviations. The data often contains extensive specialized terminology, abbreviations, and complex medical logic, requiring high semantic understanding.

Constraints on Vector Models and Indexing

The diverse sources and unstructured nature of off-label drug use data challenge the semantic understanding capabilities of vector models. For example, different literature may describe the same drug's dosage and administration with subtle variations. The model must capture and differentiate these nuances effectively. Inconsistent update frequencies necessitate incremental indexing to ensure information timeliness. Specialized terminology, abbreviations, and mixed units in documents require vector models with strong medical domain knowledge embedding capabilities to avoid recall bias due to unrecognized terms. Furthermore, data may contain classification information like evidence levels and recommendation strengths. The vector index must incorporate metadata filtering during recall to ensure response accuracy and reliability. Understanding complex medical logic also requires vector models to effectively encode contextual information to handle loosely associated but semantically relevant queries.

Configuration Guidelines

Configuration ItemRecommended ValueRationale
embedding_modeltext-embedding-ada-002 or Doubao-embeddingProvides good semantic understanding in general and medical domains, supports multiple languages.
Chunk size (Chunk Length)500–800 characters (characters)Balances semantic completeness with vector model input length limits, reducing long text truncation risk.
Chunk Overlap Length (Chunk Overlap Length)50–100 characters (characters)Ensures contextual continuity, preventing critical information from being cut off by chunk boundaries.
Recall count (Recall Count)8–12 entries (items)Balances retrieval efficiency with subsequent re-ranking processing load, while ensuring recall relevance.
Similarity threshold (Similarity Threshold)0.75–0.85Adjust based on actual testing to ensure recall results are relevant but not overly broad.
Rerank result count (Re-ranked Return Count)3–5 entries (items)Further refines recall results, improving the precision of the final response.

Common Pitfalls

  • The knowledge base index status shows "Not Ready," preventing expected query results. This indicates an abnormal interruption during document chunking or vector embedding. Check logs for error messages.
  • User queries return irrelevant responses or "No relevant information found" messages. This may be due to the vector model failing to effectively understand query semantics or overly large knowledge base document chunk granularity, leading to relevant information not being recalled.
  • Responses cite inaccurate dosage and administration or adverse reactions. This often occurs when documents contain multiple versions or conflicting information, and the vector index fails to effectively filter using metadata or evidence levels.

Validation Steps

  • Query with a batch of test questions containing key information such as drug names, indications, and dosage and administration. Check if the recalled results include the expected document snippets and compare the semantic consistency between the recalled text and the original document.
  • Adjust the Similarity threshold (Similarity Threshold) and observe changes in recall count and relevance until a balance between recall rate and accuracy is achieved.
  • Perform incremental updates for documents with different update frequencies in the knowledge base. Verify the retrieval effectiveness of both new and old documents to ensure the update mechanism functions correctly.
  • Simulate queries involving specialized medical terminology and abbreviations. Check if the vector model correctly understands and recalls relevant documents.

Note: The values provided are common starting points. Measure against specific samples to determine optimal configurations.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.