Vector Models and Indexing for Rare Disease Quality Documents

Rare disease quality documents originate from various sources. These include clinical trial reports from pharmaceutical companies, drug inserts

Data Characteristics for This Category

Rare disease quality documents originate from various sources. These include clinical trial reports from pharmaceutical companies, drug inserts, regulatory review guidelines, academic research papers, and patient organization records. Document update frequencies vary; drug inserts and regulatory guidelines might be revised annually, while clinical research data updates continuously with trial progress. Document structures typically contain extensive medical terminology, abbreviations, gene locus information, dosage units (e.g., mg/kg, IU), and complex tabular data. Fields cover gene mutation types, disease phenotypes, treatment plans, and adverse reactions. Descriptions are often highly specialized and context-dependent.

Constraints Imposed on Vector Models and Indexing

The specialized and terminology-dense nature of rare disease documents requires vector models to accurately capture the deep semantics of medical concepts and differentiate subtle disease phenotype variations. The dynamic nature of document updates demands real-time and incremental update capabilities from the index to ensure recall result timeliness. Complex document structures, especially nested tables and charts, mean traditional text segmentation methods might lose important associations, necessitating more refined preprocessing strategies. Additionally, precise matching of special fields like gene loci and dosage units challenges the vector model's encoding ability and index query accuracy. This prevents critical information loss due to semantic drift. Heterogeneous data from multiple sources also increases the difficulty of vector space alignment.

Configuration Strategy

Configuration ItemRecommended ValueRationale
embedding_modeltext-embedding-ada-002 or bce-embedding-base_v1Balances semantic understanding capability with cost, offering good support for medical terminology.
Chunk size (Segment Length)500–700 characters (characters)Balances context completeness with vector model processing efficiency, preventing key information dilution by overly long segments.
Chunk Overlap Length (Segment Overlap Length)100 characters (characters)Ensures contextual continuity at segment boundaries, reducing the risk of critical information truncation.
Similarity threshold (Similarity Threshold)0.78Ensures highly relevant recall results for rare disease queries, filtering out low-quality matches.
Recall count (Recall Count)Top 10 entries (top 10)Provides enough relevant document snippets for subsequent large language model comprehensive analysis, while avoiding information overload.
Rerank result count (Rerank Return Count)Top 3 entries (top 3)Further refines results, prioritizing the most relevant snippets to enhance user experience.

Common Pitfalls

  • An error message stating "This Token Is Not Authorized To Use The Model:text-embedding-3-large" (This token is not authorized to use model: text-embedding-3-large) typically indicates the configured API Key is not enabled or lacks permissions for the specific Embedding model.
  • Long knowledge base search response times can result from using a locally deployed vector model with high computational resource demands (e.g., shaw/dmeta-embedding-zh) on insufficient hardware (e.g., 8Core64gMemory,Graphics CardRTX2070 - 8 cores, 64GB RAM, RTX2070 graphics card), or from suboptimal index optimization leading to inefficient queries.
  • The vector model may fail to correctly identify specific gene loci or dosage units in rare disease documents, leading to semantic deviations in recall results. This occurs because the model's training data lacks sufficient medical domain-specific corpus.

Configuration Verification

  • Upload a batch of documents containing rare disease-specific gene sequences and drug dosage information. Query using this information and check if recall results include precisely matched original text snippets.
  • For rare disease questions with known answers, conduct multiple rounds of questioning to evaluate the accuracy and completeness of recall results. Adjust the Similarity threshold (Similarity Threshold) based on evaluation outcomes.
  • Monitor the average response time for knowledge base queries to ensure query latency remains within acceptable limits under different loads.
  • Check FastGPT system logs to confirm successful Embedding model calls and the absence of error codes due to API Key or channel configuration issues.

The values provided are common starting points and should be measured against specific samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.