Vector Models and Indexing for Procurement Listing Products

Procurement listing product data in the biopharmaceutical industry originates from provincial drug and medical device centralized procurement

Data Characteristics for this Product Category

Procurement listing product data in the biopharmaceutical industry originates from provincial drug and medical device centralized procurement platforms, as well as internal hospital procurement systems. This data updates frequently, typically weekly or monthly, with new product listings, price adjustments, or status changes. The document structure is primarily structured or semi-structured, such as Excel tables, XML files, or web scraping results. Core fields include product name, generic name, manufacturer, specifications, listed price, registration certificate number, medical insurance code, procurement batch, and expiration date. Price fields often contain multiple units like "RMB/box," "RMB/piece," or "RMB/tablet." Specification fields frequently include complex mixed English and Chinese descriptions and special characters.

Constraints Imposed by these Characteristics on Vector Models and Indexing

High update frequency requires the vector index to support efficient incremental updates, avoiding lengthy full rebuilds. The mix of structured and semi-structured data challenges text preprocessing and field selection, necessitating accurate extraction of key information for vectorization while handling noise in unstructured descriptions. The complexity and diversity of fields like product name and specifications mean that purely lexical matching yields poor recall; vector models must capture semantic similarity. Multi-unit price fields require unit normalization or intelligent matching during queries to ensure accurate price consultations. Exact match fields like registration certificate numbers and medical insurance codes need separate handling and should not rely solely on vector similarity.

Configuration Guidelines

Configuration ItemRecommended ValueRationale
Chunk Length300–500 charactersEnsures critical information stays within a single chunk, balancing context completeness and vector model processing capacity
Overlap Length50 charactersIncreases contextual continuity, improving recall accuracy for information spanning chunks
embedding_modeltext-embedding-ada-002 or local modelBalances model performance and deployment cost; local models can include bge-large-zh-v1.5
Vector Database TypePostgreSQL + pgvectorBalances ease of use, community support, and good compatibility with structured data, supporting efficient approximate nearest neighbor search
Recall Count10–20 itemsRecalls a sufficient number of relevant document chunks as initial candidates for subsequent reranking
Similarity Threshold0.7–0.8Filters results highly semantically relevant to the query, reducing noise; specific values require calibration against actual data and models

Three Common Pitfalls

  • The knowledge base remains in an indexing state for an extended period, displaying "Indexing..." This can be due to document parsing timeouts or vector model call failures leading to task accumulation.
  • Query results show low matching accuracy for product names and specifications, displaying irrelevant product information. This may be caused by insufficient text preprocessing, failing to effectively extract core fields for vectorization.
  • API calls return an HTTP 503 error code, indicating "current group default for model text-embedding." This typically means the configured embedding_model interface service is unavailable or access permissions are problematic.

How to Verify Configuration

  • Upload a procurement listing document containing multiple product entries. Check if the knowledge base status eventually displays "Indexing Complete" and if indexing time is within acceptable limits.
  • Use a query statement with different product names, specifications, and prices. Observe if core fields like product name and listed price in the recalled results are accurate and highly relevant to the query.
  • Through the FastGPT backend's "Data Management" or "Knowledge Base Details" page, check if the number of vector indexes roughly matches the number of uploaded document chunks. Review the embedding_model call logs for any abnormal errors.

Note: The values provided are common starting points and should be measured against the reader's own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.