Knowledge Base Retrieval and Recall for Indication-Based Medication Q&A

Indication data primarily originates from drug inserts, national drug administration databases, medical guidelines, and clinical research reports.

Data Characteristics

Indication data primarily originates from drug inserts, national drug administration databases, medical guidelines, and clinical research reports. Update frequency is relatively stable, typically aligning with drug approval and guideline revision cycles (quarterly or annually). Urgent drug information updates may occur more frequently. Document structures are often semi-structured, including fields such as generic name, brand name, indication description, dosage and administration, contraindications, and adverse reactions. Indication descriptions are usually natural language text, potentially containing disease classification codes (e.g., ICD-10), symptom descriptions, and disease stages. Field values lack a unified unit; indications are descriptive text.

Constraints on Knowledge Base Retrieval and Recall

The semi-structured nature of indication data limits the accuracy of pure keyword matching recall, requiring semantic understanding. A stable update frequency means knowledge base construction does not require extreme real-time capabilities, but regular full or incremental updates are necessary to maintain data timeliness. Documents contain extensive medical terminology and disease descriptions, demanding strong domain-specific vocabulary understanding from the retrieval model to avoid recall omissions due to synonyms, near-synonyms, or abbreviations. Indication descriptions vary in length, from a single disease name to detailed pathophysiological processes. This affects text segmentation strategies and context window size settings. The lack of unified units for fields means numerical range filtering is not possible; recall primarily relies on text similarity.

Configuration Settings

Configuration ItemRecommended ValueRationale
Chunk size (Chunk Length)500–800 charactersBalances information completeness and retrieval efficiency for moderately long indication descriptions.
Chunk Overlap Length (Chunk Overlap Length)80 charactersEnsures contextual continuity and prevents critical information from being split.
Recall count (Number of Retrieved Chunks)Top 5Balances recall breadth with the computational cost of subsequent reranking.
Similarity threshold (Similarity Threshold)Calibrated by actual measurementAdjusted based on the business trade-off between precision and recall, using a test set.
Rerank result count (Number of Reranked Results)Top 3Further filters for the most relevant results, reducing the burden on the large language model.
maxContext2000–3000 tokensEnsures the large language model receives sufficient context to process complex indication descriptions.

Common Mistakes

  • Slow response or timeout errors after clicking on the knowledge base may be due to overly fine-grained knowledge base chunking or excessive data volume, leading to increased vector retrieval pressure.
  • Retrieval results containing numerous irrelevant indication details, where similarity scores are high but semantics do not match, indicate a lack of semantic understanding of domain-specific vocabulary. Generic vector models fail to accurately capture medical concepts.
  • New indication information not being retrieved after a knowledge base update (e.g., querying for the latest drugs returns old information) may be due to untimely knowledge base index rebuilding or flaws in the incremental update mechanism.

Validation Steps

  • Select a batch of indication queries containing both new and old knowledge. Check if retrieval results include all relevant and up-to-date information, and evaluate their ranking priority.
  • For a set of test questions containing medical terms, abbreviations, and synonyms, verify if the knowledge base can accurately recall corresponding indication descriptions. Evaluate the recall rate.
  • Conduct simulated concurrent queries during peak hours. Observe if knowledge base response times remain within an acceptable range and check for error logs.

The values provided are common starting points and should be measured against specific samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.