Data Characteristics for this Category
Pharmacoeconomics data primarily originates from clinical trial reports, real-world evidence (RWE) data, drug registration approvals, national medical insurance catalogs, drug procurement price data, and relevant policies and regulations. Data update frequencies vary; clinical guidelines and medical insurance policies may update annually or irregularly, while drug price data updates more frequently. Document structures are diverse, including structured database records, semi-structured reports (e.g., HTA reports), and unstructured research papers and policy documents. Key fields include drug name, indications, treatment plans, costs (direct and indirect), effects (QALY, DALY), utility values, payer information, and research methods and results. Cost units often involve currencies like RMB and USD, while effect units are typically Quality-Adjusted Life Years (QALY).
Constraints Imposed by These Characteristics on Knowledge Base Retrieval and Recall
The diversity of pharmacoeconomics data requires the knowledge base to handle multiple document formats and effectively segment and index different data types. Due to the frequent updates of some data, the knowledge base must support incremental updates and version management to ensure the timeliness of retrieval results. The accuracy of key numerical values like costs and effects is crucial for result reliability, necessitating high-precision entity recognition and numerical extraction capabilities. Potential cross-terminology, abbreviations, and result discrepancies due to different research methods within documents demand higher semantic understanding from recall algorithms. Some document content may contain conflicts or discrepancies, requiring retrieval results to provide multi-source information and assist users in identifying potential contradictions.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale for this Value |
|---|---|---|
Chunk size (Segment Length) | 500–800 characters | Pharmacoeconomics reports often contain detailed arguments; longer segments help retain contextual integrity. |
Chunk Overlap Length (Segment Overlap Length) | 100 characters | Ensures semantic continuity between adjacent paragraphs, preventing key information from being cut off. |
Recall count (Recall Count) | Top 5–8 entries | Given the complexity of pharmacoeconomics analysis, multi-perspective information support is needed. |
Similarity threshold (Similarity Threshold) | 0.75 | Balances precise matching and semantic relevance, avoiding the omission of highly relevant documents with different phrasing. |
Rerank result count (Rerank Return Count) | 3 entries | Further focuses on the most core and relevant pieces of information after initial recall. |
Text Understanding Model | FastGPT-Rerank-v1.0 | Optimized for the Chinese biomedical domain, improving semantic understanding and recall accuracy. |
Three Common Mistakes
- The document list returned after a knowledge base query is empty, or only a few irrelevant documents are returned. This may be due to a
Similarity threshold(Similarity Threshold) set too high, strictly filtering out some relevant but insufficiently similar documents. - The AI's response cites outdated policy or pricing information instead of the latest data from the knowledge base. This usually happens because the knowledge base's update mechanism was not triggered in time, or indexing lagged behind data source updates.
- When configuring a workflow, after dynamically passing values to the
knowledgeSearchvariable, the AI's response does not cite the documents. This could be because the dynamically passed query parameters do not match the knowledge base index terms, or theRecall count(Recall Count) is set too low, failing to recall enough documents for the LLM to cite.
How to Confirm Correct Configuration
- Query typical pharmacoeconomics questions and check if the returned reference documents contain key cost, effect values, and relevant research methods.
- Select recently updated medical insurance policies or drug price data to verify whether the knowledge base query results can recall and cite this latest information.
- For potential conflicting data points (e.g., different studies' cost-effectiveness analyses for the same drug), test whether query results can simultaneously recall and present multi-source information to assist users in comparison.
- Evaluate whether the knowledge base document sources cited in the AI's response are diverse, covering clinical guidelines, HTA reports, price data, and other types.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.