Knowledge Base Retrieval and Recall for Pharmacoeconomics Products

Pharmacoeconomics data primarily originates from clinical trial reports, real-world evidence (RWE) data, drug registration approvals, national medical

Data Characteristics for this Category

Pharmacoeconomics data primarily originates from clinical trial reports, real-world evidence (RWE) data, drug registration approvals, national medical insurance catalogs, drug procurement price data, and relevant policies and regulations. Data update frequencies vary; clinical guidelines and medical insurance policies may update annually or irregularly, while drug price data updates more frequently. Document structures are diverse, including structured database records, semi-structured reports (e.g., HTA reports), and unstructured research papers and policy documents. Key fields include drug name, indications, treatment plans, costs (direct and indirect), effects (QALY, DALY), utility values, payer information, and research methods and results. Cost units often involve currencies like RMB and USD, while effect units are typically Quality-Adjusted Life Years (QALY).

Constraints Imposed by These Characteristics on Knowledge Base Retrieval and Recall

The diversity of pharmacoeconomics data requires the knowledge base to handle multiple document formats and effectively segment and index different data types. Due to the frequent updates of some data, the knowledge base must support incremental updates and version management to ensure the timeliness of retrieval results. The accuracy of key numerical values like costs and effects is crucial for result reliability, necessitating high-precision entity recognition and numerical extraction capabilities. Potential cross-terminology, abbreviations, and result discrepancies due to different research methods within documents demand higher semantic understanding from recall algorithms. Some document content may contain conflicts or discrepancies, requiring retrieval results to provide multi-source information and assist users in identifying potential contradictions.

Configuration Guidelines

Configuration ItemRecommended ValueRationale for this Value
Chunk size (Segment Length)500–800 charactersPharmacoeconomics reports often contain detailed arguments; longer segments help retain contextual integrity.
Chunk Overlap Length (Segment Overlap Length)100 charactersEnsures semantic continuity between adjacent paragraphs, preventing key information from being cut off.
Recall count (Recall Count)Top 5–8 entriesGiven the complexity of pharmacoeconomics analysis, multi-perspective information support is needed.
Similarity threshold (Similarity Threshold)0.75Balances precise matching and semantic relevance, avoiding the omission of highly relevant documents with different phrasing.
Rerank result count (Rerank Return Count)3 entriesFurther focuses on the most core and relevant pieces of information after initial recall.
Text Understanding ModelFastGPT-Rerank-v1.0Optimized for the Chinese biomedical domain, improving semantic understanding and recall accuracy.

Three Common Mistakes

  • The document list returned after a knowledge base query is empty, or only a few irrelevant documents are returned. This may be due to a Similarity threshold (Similarity Threshold) set too high, strictly filtering out some relevant but insufficiently similar documents.
  • The AI's response cites outdated policy or pricing information instead of the latest data from the knowledge base. This usually happens because the knowledge base's update mechanism was not triggered in time, or indexing lagged behind data source updates.
  • When configuring a workflow, after dynamically passing values to the knowledgeSearch variable, the AI's response does not cite the documents. This could be because the dynamically passed query parameters do not match the knowledge base index terms, or the Recall count (Recall Count) is set too low, failing to recall enough documents for the LLM to cite.

How to Confirm Correct Configuration

  • Query typical pharmacoeconomics questions and check if the returned reference documents contain key cost, effect values, and relevant research methods.
  • Select recently updated medical insurance policies or drug price data to verify whether the knowledge base query results can recall and cite this latest information.
  • For potential conflicting data points (e.g., different studies' cost-effectiveness analyses for the same drug), test whether query results can simultaneously recall and present multi-source information to assist users in comparison.
  • Evaluate whether the knowledge base document sources cited in the AI's response are diverse, covering clinical guidelines, HTA reports, price data, and other types.

The values provided are common starting points and should be measured against the reader's own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.