Data Characteristics
Pharmacoeconomics regulatory submission data comes from various sources. These include clinical trial reports, real-world study data, health technology assessment reports, model construction files, and national or regional medical insurance policy documents. Data updates are infrequent, typically occurring with new drug launches, indication expansions, or adjustments to medical insurance catalogs. Document structures are complex. They often contain numerous tables, charts, statistical data, and specialized terminology. Core fields like QALYs (Quality-Adjusted Life Years), ICER (Incremental Cost-Effectiveness Ratio), cost components, effectiveness indicators, discount rates, and sensitivity analysis parameters have clear quantitative units and professional definitions.
Constraints on Vector Models and Indexing
The specialized and complex nature of pharmacoeconomics data requires vector models to accurately capture subtle semantic differences and numerical relationships. For example, different discount rates or sensitivity analysis assumptions can significantly change ICER results. Vector models must distinguish these minor but critical parameter settings. Documents frequently contain tables and charts. This challenges traditional text segmentation methods. More refined text preprocessing and chunking strategies are necessary to ensure the integrity of tabular data and its context. Data update frequency is low. This relaxes real-time indexing requirements. However, high accuracy is needed for historical version management and retrieval to support data traceability across different submission stages.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Chunk Size) | 800–1200 characters | Balances context completeness with vector model processing capabilities. Avoids information overload in a single chunk. |
Chunk Overlap Length (Overlap Size) | 100–200 characters | Ensures semantic continuity between adjacent paragraphs, especially when processing complex logic and tabular data. |
Recall count (Recall Count) | Top 8–12 | Covers a wider range of potentially relevant information. Improves retrieval comprehensiveness. |
Similarity threshold (Similarity Threshold) | Calibrate based on actual measurements | Pharmacoeconomics concepts are highly specialized. Adjust based on specific datasets to ensure high accuracy. |
Rerank result count (Reranked Return Count) | Top 3–5 | Focuses on the most relevant results. Reduces manual filtering burden for engineers. |
PARSE_FILE_TIMEOUT_SECONDS | 300 seconds | Addresses the time required for parsing large PDF files and complex tables. |
Common Pitfalls
- Symptom: Retrieval results contain many irrelevant or duplicate paragraphs, and core data is missing. Reason: The chunking strategy did not effectively handle tables or charts. This led to critical information being truncated or context lost.
- Symptom: The configured model does not appear in the text understanding model dropdown when creating a knowledge base. Reason: The model channel configuration is correct, but the model is not registered at the system level or its availability status is abnormal.
- Symptom: Retrieval accuracy is much lower than expected after batch uploading documents and vectorizing them. Reason: Lack of preprocessing for pharmacoeconomics specialized terminology and numerical units. This results in insufficient understanding of key information by the model.
Verification Steps
- Upload documents containing typical pharmacoeconomics data (e.g., ICER tables, descriptions of sensitivity analysis charts). Check the vectorized chunk content. Ensure tabular data and key numerical values are fully captured.
- Perform a series of retrievals with pharmacoeconomics specialized terms and complex queries. Evaluate the relevance and accuracy of the recalled results. Verify that returned chunks contain the core information required by the query.
- For different queries, observe the impact of adjusting
Similarity threshold(Similarity Threshold) andRecall count(Recall Count) on the result list. Determine a suitable threshold range for the business scenario.
The values provided are common starting points. Measure them against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.