Vector Models and Indexing for Tender Procurement Drug Vigilance

Tender procurement drug vigilance data originates from provincial drug procurement platforms, healthcare security bureau websites, and drug

Data Characteristics

Tender procurement drug vigilance data originates from provincial drug procurement platforms, healthcare security bureau websites, and drug administration announcements. Data updates frequently, typically monthly or quarterly. Documents come in various formats, including PDF for tender announcements, winning bids, and listing notices, and Excel or XML for drug catalogs, price lists, and adverse reaction monitoring reports. Document structures are relatively fixed. For example, tender announcements include fields like drug name, specification, manufacturer, and declared price. Adverse reaction reports include event descriptions, drug batch numbers, and patient information. Fields often involve specialized terms such as generic drug names, brand names, dosages, specifications, and batch numbers. Units are primarily milligrams, milliliters, tablets, vials, and boxes, with multiple representation forms.

Constraints on Vector Models and Indexing

The diversity and update frequency of tender procurement data impose specific requirements on vector models and indexing. First, multiple document formats like PDF and Excel necessitate a unified document parsing capability to ensure complete information extraction. Second, frequent data updates mean the knowledge base requires efficient incremental update mechanisms to avoid resource consumption from full rebuilds. Documents contain a large amount of structured and semi-structured data, challenging the semantic understanding capabilities of vector models. Models must distinguish key information such as drug names, specifications, and prices. Additionally, the standardization of specialized terms and units requires vector models to possess domain-specific knowledge to accurately capture drug name variations or synonyms. Indexing strategies must balance recall accuracy and efficiency, especially when processing large volumes of similar drug information, to avoid confusion.

Configuration Guidelines

Configuration ItemSuggested ValueRationale
Chunk size (Segment Length)800–1200 characters (characters)Balances semantic completeness with query efficiency, prevents excessively long segments from diluting key information
Chunk Overlap Length (Segment Overlap Length)100–200 characters (characters)Ensures semantic continuity across segments, improves recall quality
Vector Model (Vector Model)bge-large-zh-1.5 or compatible modelOptimized for Chinese biomedical domain, good semantic understanding capabilities
Recall count (Recall Count)Top 8–12 entries (top 8–12 items)Balances recall breadth with efficiency of subsequent re-ranking
Similarity threshold (Similarity Threshold)Calibrated by actual measurementEnsures relevant results are recalled, avoids interference from irrelevant information
PARSE_FILE_TIMEOUT_SECONDS600 seconds (seconds)Addresses potentially long parsing times for large PDF or Excel files

Common Pitfalls

  • Knowledge base document status remains "indexing" for extended periods because parsing large or complex tender procurement files exceeds the default parsing timeout.
  • Inaccurate query results after uploading locally pre-segmented documents to the server because different vector models are used locally and on the server, leading to inconsistent vector spaces.
  • Different specifications or batch numbers of the same drug are confused in retrieval results because the vector model lacks sufficient semantic understanding of specialized fields, or the indexing strategy fails to effectively distinguish subtle differences.

Verification Steps

  • Upload typical tender announcements and adverse reaction reports. Check if all documents are successfully parsed and indexed.
  • Perform queries for specific drug names, specifications, and manufacturers. Verify if recall results accurately include relevant information and evaluate the recall count.
  • Simulate adverse reaction event queries. Verify if the system can associate with the correct drug batch numbers and event descriptions, and evaluate the similarity threshold.
  • Observe query performance after incremental knowledge base updates. Ensure newly uploaded data is retrievable in a timely manner.

Note: The values provided are common starting points. They should be measured against specific samples to determine optimal configurations.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.