Data Characteristics for This Category
Pharmacoeconomics regulations and standard documents originate from policy papers and guidelines published by the National Health Commission, the National Healthcare Security Administration, and provincial/municipal healthcare security bureaus, as well as expert consensuses from industry associations. These documents typically come in PDF, Word, or structured text formats. Update frequency is relatively low, often quarterly or annually. Documents have a rigorous structure, containing extensive specialized terminology, abbreviations, formulas, and tabular data. Examples include drug registration prices, medical insurance payment standards, cost-effectiveness analysis model parameters, and Quality-Adjusted Life Year (QALY) calculation methods. Fields cover drug names, indications, reimbursement scope, payment restrictions, price composition, clinical efficacy data, and adverse event rates, often with clear units of measurement.
Constraints on Knowledge Base Retrieval and Recall from These Characteristics
Specialized terminology and abbreviations in pharmacoeconomics documents require high precision in word segmentation and semantic understanding for the knowledge base. This avoids recall bias due to lexical ambiguity. Complex document structures, including numerous tables and formulas, mean that simple text segmentation can split critical information, affecting retrieval completeness. For instance, a drug's medical insurance payment standard might be spread across multiple paragraphs or even tables, requiring intelligent identification and integration. Lower update frequency allows for a more relaxed knowledge base indexing update strategy, but each update needs to ensure traceability and management of historical versions. Furthermore, many internal Confluence pages or specific document formats require the knowledge base to have robust heterogeneous document parsing capabilities to convert non-standard formats into retrievable text.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Chunk Size) | 800–1200 characters (characters) | Balances paragraph completeness and retrieval efficiency, preventing truncation of key descriptions or table content. |
Chunk Overlap Length (Chunk Overlap) | 100–200 characters (characters) | Ensures contextual continuity, especially where professional concepts or logical arguments are linked. |
Recall count (Recall Count) | 8 entries (items) | Covers more potentially relevant information, addressing specialized terminology and multi-dimensional queries. |
Similarity threshold (Similarity Threshold) | 0.75 | Guarantees precision of recall results, filtering out low-relevance policy provisions. |
Rerank result count (Rerank Return Count) | 4 entries (items) | Focuses on the most core and direct answers, improving final response quality. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds (seconds) | Handles parsing of large policy documents, preventing content loss due to timeouts. |
Three Common Pitfalls
- Query results are empty or fail to recall expected policy documents. This happens when the knowledge base cannot correctly parse tables or specific formatted content within internal Confluence pages, leading to incomplete indexing.
- Recalled policy provisions are too broad and lack specificity. This typically occurs when the
Similarity threshold(Similarity Threshold) is set too low, recalling a large volume of low-relevance text. - Tool calls to online search return empty responses. This might be due to system egress network configuration restrictions, preventing FastGPT from accessing external network resources for auxiliary retrieval.
How to Verify Configuration
- Select typical pharmacoeconomics questions and verify if the knowledge base can recall at least one policy provision containing core keywords.
- For policy documents with tables or formulas, check if recall results accurately present relevant data or calculation methods.
- Test different query methods (e.g., abbreviations, full names, conceptual descriptions) and observe the stability and consistency of recall results.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.