Citing and Tracing Sources in Pharmacoeconomics Quality Documents

Pharmacoeconomics quality documents primarily include cost-effectiveness analysis reports, budget impact analysis reports, drug value assessment

Data Characteristics in This Category

Pharmacoeconomics quality documents primarily include cost-effectiveness analysis reports, budget impact analysis reports, drug value assessment reports, and related methodological guidelines and consensuses. These documents typically originate from pharmaceutical companies, academic institutions, government health departments, and specialized consulting firms. Their update frequency is influenced by factors such as new drug launches, revisions to treatment guidelines, and adjustments to health insurance policies. Major revisions usually occur annually or biennially, while some data (e.g., drug prices, disease prevalence) may update more frequently. Documents feature a rigorous structure, including standard sections like abstract, background, methods, results, discussion, and conclusion. Data fields cover drug prices, disease incidence, treatment pathways, clinical efficacy data (e.g., QALY, DALY), healthcare resource consumption, and societal costs. Units are diverse, such as CNY, USD, years, persons, and percentages, often accompanied by sensitivity analysis results and uncertainty intervals.

Constraints Imposed by These Characteristics on "Citing and Tracing Sources"

The diverse data sources and varying update frequencies of pharmacoeconomics documents require RAG retrieval systems to precisely pinpoint original data sources when citing and to identify differences between versions. These documents contain numerous tables, charts, and statistical data, which can lead to loss of context with traditional text segmentation, affecting citation accuracy. For example, a QALY value's specific calculation method, assumed parameters, and uncertainty range might be distributed across different parts of a document. Furthermore, parameters and results from different economic models (e.g., decision trees, Markov models) are interrelated, requiring the ability to integrate cross-sectional information during tracing. Accurate identification of financial and medical units is also crucial to ensure unambiguous citations. The system must support range retrieval and multi-dimensional filtering for numerical data to meet the citation requirements of sensitivity analysis results.

Configuration Settings

Configuration ItemSuggested ValueRationale for Value
Chunk size500–800 charactersPharmacoeconomics document paragraphs are long, containing complex logic and data; this length helps maintain contextual integrity.
Recall countTop 8–12 entriesEnsures coverage of different data sources and model assumptions, improving comprehensive understanding of complex economic analyses.
Similarity threshold0.78–0.85Balances recall and precision, avoiding retrieval of irrelevant general medical or statistical texts.
Rerank result countTop 5 entriesPrioritizes displaying the most relevant core conclusions and key data to reduce the model's processing burden.
maxContext3000–4000 tokensAccommodates complex economic models and multi-parameter analyses, ensuring the model can process sufficient citation information.
Max Knowledge Base References3–5 PlacesAvoids model confusion from excessive citations, focusing on key evidence and data sources.

Three Common Mistakes

  • The model fails to provide specific numerical sources or report names in its answers. The symptom is an answer with only conclusions, lacking annotations that identify the report and page. This occurs because the knowledge base chunking granularity is too large or citation source information is not effectively embedded in the metadata.
  • When generating answers about cost-effectiveness, different currencies or time points of data are mixed without explanation. The symptom is inconsistent units or contradictory numerical values in the answer. This happens when numerical fields are not standardized during document processing or context recognition is insufficient.
  • The API interface returns chaotic citation information or includes a large amount of non-critical citation text. The symptom is redundant and difficult-to-read citation content. This is caused by improper settings for Max Knowledge Base References or maxContext parameters, leading the model to cite too many non-core fragments.

How to Confirm Proper Configuration

  • For typical pharmacoeconomics queries, check whether the citation sources provided in the model's answers accurately point to specific sections or page numbers of the original documents.
  • Verify that for queries involving numerical data, the model's citation results reflect the original data's units, time, and statistical scope.
  • Test pharmacoeconomics questions of varying complexity to evaluate the accuracy of generated answers and cross-reference key conclusions against the original documents.

The values provided are common starting points and should be measured against the reader's own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.