Data Characteristics
Core data for pharmacoeconomic regulatory submissions originates from clinical trial reports, real-world evidence (RWE) studies, published medical literature, government-issued drug pricing and reimbursement policy documents, and various databases (e.g., medical insurance catalogs, national centralized drug procurement data). Data update frequencies vary; policy documents may be issued or revised annually, while clinical trial data updates as research progresses. Documents are typically in PDF, Word, or Excel formats, containing extensive textual descriptions, charts, and statistical data. Key fields include drug costs, efficacy data, safety indicators, quality of life scores, cost-effectiveness ratios (ICER), and budget impact analysis results. Units involve currency (e.g., RMB yuan), time (e.g., years, months), percentages, and QALY (Quality-Adjusted Life Year).
Constraints Imposed by These Characteristics on "Citing Sources and Traceability"
The diverse and unstructured nature of pharmacoeconomic data imposes specific requirements on source citation and traceability. For example, the timeliness of policy documents requires the system to quickly identify and cite the latest version, avoiding outdated information. Complex charts and statistical data in clinical trial reports require the RAG system to accurately extract key values and conclusions, linking them to specific page numbers in the original report. Varying units and calculation logic in cost-effectiveness analysis reports demand a traceability mechanism that points to original data sources and methodology descriptions. Additionally, due to sensitive commercial and policy information, high precision and verifiability of cited content are crucial to ensure the authenticity and regulatory compliance of generated content, preventing compliance risks in submission documents due to citation errors.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk Length | 800–1200 characters | Pharmacoeconomic documents often contain lengthy descriptive text and detailed data. Longer chunks help maintain contextual integrity. |
Recall Count | Top 5–8 items | Ensures coverage of multiple relevant paragraphs from key sections such as cost-effectiveness analysis and budget impact analysis. |
Similarity Threshold | 0.75–0.82 | Balances high relevance with avoiding omission of crucial information due to differing phrasing. |
Rerank Return Count | Top 3 items | Selects the most relevant citations, reducing redundancy and improving the accuracy and conciseness of generated content. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Provides sufficient parsing time when processing large PDF documents or Excel files containing complex tables. |
Knowledge Base Q&A Max Length | 500 characters | Accommodates longer professional terms and data descriptions common in pharmacoeconomics, ensuring Q&A pair completeness. |
Common Mistakes
- Symptom: Generated submission documents lack citations for key data, or citations point to original content that does not match. Reason: Knowledge base document chunking granularity is too large, preventing the RAG model from precisely matching the smallest effective information units during recall.
- Symptom: The system fails to generate effective Q&A pairs when processing pharmacoeconomic reports with numerous charts and complex tables. Reason: The file parser has insufficient capability to extract non-text content (e.g., tabular data embedded in images) or lacks specific processing strategies for tabular data.
- Symptom: Even with a high similarity threshold configured, the system still cites outdated policy documents. Reason: The knowledge base lacks version management or timestamping for documents, preventing retrieval from prioritizing the latest version of policy documents.
Validation Steps
- Select multiple typical pharmacoeconomic reports and verify that the system-generated citations accurately point to specific paragraphs or data tables within the reports.
- Test policy documents from different time points, checking if the system citations prioritize and present the latest published versions.
- For documents containing complex cost-effectiveness analysis charts, verify that the system can accurately extract key values from the charts and trace them back to the original charts.
- Randomly select citations from generated submission documents and manually cross-reference the cited content with the original text, evaluating its accuracy within the pharmacoeconomic context.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.