Data Characteristics
Pharmacoeconomics research data originates from clinical trial reports, real-world study data, Health Technology Assessment (HTA) reports, cost-effectiveness analysis models, pharmaceutical company market access documents, and government healthcare policy files. These documents are updated infrequently, typically with new drug approvals, expanded indications, or changes in healthcare catalogs. Document structures are primarily structured or semi-structured text, containing numerous charts, tables, and specialized terminology. Fields include drug names, indications, treatment regimens, efficacy metrics (e.g., QALY, LYG), cost components (direct, indirect), effect values, and Incremental Cost-Effectiveness Ratios (ICER). Units encompass monetary units (USD, EUR, RMB), time units (years, months), quality of life units (QALY), and various statistical indicators.
Constraints Imposed by Data Characteristics on Vector Models and Indexing
The low update frequency of pharmacoeconomics documents means knowledge base reconstruction or incremental updates are not required often. The specialized terminology, abbreviations, and co-existence of multiple languages (e.g., original English and Chinese translations) in documents demand vector models with high semantic understanding to capture deep lexical relationships. The large volume of numerical data, tables, and charts challenges indexing segmentation strategies. Numerical context must be preserved. For example, an ICER value must be indexed with its corresponding cost, effect, and comparison scheme to provide meaningful answers during retrieval. Furthermore, multi-dimensional field information, such as pharmacoeconomics data from different countries or regions, requires indexing to support fine-grained metadata filtering to prevent irrelevant retrievals.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Segment Length) | 500–800 characters | Balances contextual completeness with retrieval accuracy, accommodating the length of specialized terms. |
Chunk Overlap Length (Segment Overlap Length) | 100–150 characters | Ensures semantic continuity between segments, preventing critical information from being split. |
Vector Model (Vector Model) | Doubao-embedding or text-embedding-ada-002 | Selects models with strong semantic understanding for specialized domain texts. |
Recall count (Number of Retrieved Items) | 8–12 items | Increases the probability of retrieving relevant information, covering different aspects. |
Similarity threshold (Similarity Threshold) | 0.75–0.85 | Filters irrelevant segments, ensuring the precision of retrieval results. |
Rerank result count (Number of Reranked Items) | 3–5 items | Optimizes the final presentation, focusing on the most core relevant content. |
Common Pitfalls
- Knowledge base vector model switching stalls, with no forced switch option. This occurs because model switching may require clearing and rebuilding all vector indexes. If the data volume is large or resources are insufficient, this process is time-consuming, leading to delayed frontend response or backend task blockage.
- When configuring locally or privately deployed embedding models like
Doubao-embedding, test reports show a404 page not founderror. This is typically due to incorrect API address configuration or the FastGPT deployment environment being unable to access the embedding service port. - When indexing image or table content, retrieved results lack critical numerical information, such as ICER values or cost data. This happens because the default text segmentation strategy does not effectively process non-text content, leading to numerical values being separated from their context, or image content not being correctly OCR'd and converted into indexable text.
Validation Steps
- Upload typical pharmacoeconomics documents (e.g., an HTA report). Check the knowledge base indexing status to ensure all files are indexed.
- Conduct Q&A tests using specialized terminology and specific cost-effectiveness ratios from the documents. Verify that retrieved results include relevant segments and numerical values.
- Attempt to query pharmacoeconomics data from different countries or regions. Observe whether the metadata filtering function effectively filters out corresponding information.
- Regularly check logs to ensure no abnormal errors in vector model calls and that all indexing tasks complete successfully.
The values provided are common starting points. Measure them against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.