Knowledge Base Retrieval and Recall for Tender Bidding and Listing Pharmacovigilance

Tender bidding and listing pharmacovigilance data originates from drug centralized procurement platforms, official medical insurance bureau websites

Data Characteristics

Tender bidding and listing pharmacovigilance data originates from drug centralized procurement platforms, official medical insurance bureau websites, and tender announcements from drug regulatory authorities. This data updates frequently, typically quarterly or annually in bulk. Emergency updates occur when policies change or unforeseen events happen. Document structures are primarily tabular, often published as PDFs or Excels. Fields include drug name, generic name, manufacturer, dosage form, specification, price, procurement cycle, adverse reaction monitoring requirements, and risk control plans. Adverse reaction monitoring requirements and risk control plans often embed as unstructured text, requiring additional parsing. The price field frequently involves multiple unit conversions, such as "yuan/box" and "yuan/tablet," which require standardization.

Constraints on Knowledge Base Retrieval and Recall

The tabular nature of tender bidding data requires the knowledge base to recognize semantic relationships between rows and columns during chunking. This avoids isolating single-line text and losing context. High-frequency updates mean the knowledge base needs efficient incremental update mechanisms to reduce duplicate indexing overhead. Embedded unstructured text demands stronger semantic understanding from the model to precisely extract pharmacovigilance-related key information from lengthy descriptions. Multi-unit price fields require standardization before vectorization to ensure correct matching of price information across different units during retrieval. Additionally, tender bidding retrieval often combines precise matching with fuzzy semantic retrieval. For example, it needs to precisely find the latest listed price of a specific drug in a certain province, and also fuzzily match adverse reaction risk descriptions for similar drugs.

Configuration Settings

Configuration ItemSuggested ValueRationale
chunk_size500–800 charactersBalances context integrity for tabular rows and unstructured text, avoiding overly long or short chunks
chunk_overlap50 charactersPreserves contextual continuity, ensuring semantic coherence across chunks
recall_counttop 8Covers potentially relevant information, balancing recall rate and computational cost
similarity_thresholdCalibrate by actual measurementAdjust through iterative testing based on actual retrieval performance and business requirements
rerank_counttop 5Focuses on the most relevant results, improving the precision of the final presentation
maxContext3000 TokensSupports full inclusion of longer tender announcement originals or adverse reaction descriptions

Common Pitfalls

  • After document upload, the page refresh does not complete, preventing further operations. This typically results from file parsing timeouts, especially when handling large PDF or Excel files, due to a small PARSE_FILE_TIMEOUT_SECONDS parameter.
  • The knowledge base fails to recall precise information related to tender bidding prices, even when explicitly present in the text. This may occur if the chunking strategy is too aggressive, separating critical numbers and their units in tables into different chunks, leading to semantic loss.
  • Semantic retrieval cannot find clearly relevant pharmacovigilance descriptions, but full-text search can. This indicates that the chosen embedding model lacks sufficient understanding of specific industry terminology, or the vector library indexing strategy is unsuitable for such specialized texts.

How to Confirm Correct Configuration

  • Upload typical tender announcement PDF or Excel files. Check if knowledge base chunking is reasonable and if tabular row data maintains complete semantics.
  • Construct precise queries including drug generic name, specification, and price unit. Verify if the correct listed information is recalled and check the completeness of the recalled content.
  • For a drug's adverse reaction description, construct multiple semantically similar but differently worded queries. Observe the ranking and relevance of recall results, and set an appropriate similarity_threshold based on expert feedback.
  • Simulate tender bidding data updates. Verify if the knowledge base's incremental update function works correctly and if newly uploaded data is indexed and retrieved promptly.

Note: The values provided are common starting points. Measure them against your own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.