Data Characteristics
Tender bidding and listing pharmacovigilance data originates from drug centralized procurement platforms, official medical insurance bureau websites, and tender announcements from drug regulatory authorities. This data updates frequently, typically quarterly or annually in bulk. Emergency updates occur when policies change or unforeseen events happen. Document structures are primarily tabular, often published as PDFs or Excels. Fields include drug name, generic name, manufacturer, dosage form, specification, price, procurement cycle, adverse reaction monitoring requirements, and risk control plans. Adverse reaction monitoring requirements and risk control plans often embed as unstructured text, requiring additional parsing. The price field frequently involves multiple unit conversions, such as "yuan/box" and "yuan/tablet," which require standardization.
Constraints on Knowledge Base Retrieval and Recall
The tabular nature of tender bidding data requires the knowledge base to recognize semantic relationships between rows and columns during chunking. This avoids isolating single-line text and losing context. High-frequency updates mean the knowledge base needs efficient incremental update mechanisms to reduce duplicate indexing overhead. Embedded unstructured text demands stronger semantic understanding from the model to precisely extract pharmacovigilance-related key information from lengthy descriptions. Multi-unit price fields require standardization before vectorization to ensure correct matching of price information across different units during retrieval. Additionally, tender bidding retrieval often combines precise matching with fuzzy semantic retrieval. For example, it needs to precisely find the latest listed price of a specific drug in a certain province, and also fuzzily match adverse reaction risk descriptions for similar drugs.
Configuration Settings
| Configuration Item | Suggested Value | Rationale |
|---|---|---|
chunk_size | 500–800 characters | Balances context integrity for tabular rows and unstructured text, avoiding overly long or short chunks |
chunk_overlap | 50 characters | Preserves contextual continuity, ensuring semantic coherence across chunks |
recall_count | top 8 | Covers potentially relevant information, balancing recall rate and computational cost |
similarity_threshold | Calibrate by actual measurement | Adjust through iterative testing based on actual retrieval performance and business requirements |
rerank_count | top 5 | Focuses on the most relevant results, improving the precision of the final presentation |
maxContext | 3000 Tokens | Supports full inclusion of longer tender announcement originals or adverse reaction descriptions |
Common Pitfalls
- After document upload, the page refresh does not complete, preventing further operations. This typically results from file parsing timeouts, especially when handling large PDF or Excel files, due to a small
PARSE_FILE_TIMEOUT_SECONDSparameter. - The knowledge base fails to recall precise information related to tender bidding prices, even when explicitly present in the text. This may occur if the chunking strategy is too aggressive, separating critical numbers and their units in tables into different chunks, leading to semantic loss.
- Semantic retrieval cannot find clearly relevant pharmacovigilance descriptions, but full-text search can. This indicates that the chosen embedding model lacks sufficient understanding of specific industry terminology, or the vector library indexing strategy is unsuitable for such specialized texts.
How to Confirm Correct Configuration
- Upload typical tender announcement PDF or Excel files. Check if knowledge base chunking is reasonable and if tabular row data maintains complete semantics.
- Construct precise queries including drug generic name, specification, and price unit. Verify if the correct listed information is recalled and check the completeness of the recalled content.
- For a drug's adverse reaction description, construct multiple semantically similar but differently worded queries. Observe the ranking and relevance of recall results, and set an appropriate
similarity_thresholdbased on expert feedback. - Simulate tender bidding data updates. Verify if the knowledge base's incremental update function works correctly and if newly uploaded data is indexed and retrieved promptly.
Note: The values provided are common starting points. Measure them against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.