Data Characteristics
Bidding and listing regulation data typically originates from public resource trading centers, medical insurance bureaus, and internal compliance departments of pharmaceutical companies across various provinces and cities. Updates are driven by policy adjustments and annual procurement cycles, occurring quarterly or annually. Urgent notices are published immediately. Documents are primarily in PDF, Word, and HTML formats. Content includes original policies, interpretation documents, operational guidelines, and FAQs. Common fields include policy number, publication date, implementation date, scope, specific clauses, drug or device classification, application conditions, and evaluation criteria. Units are often dates, text descriptions, numerical ranges, or percentages.
Constraints on Knowledge Base Retrieval and Recall
The diverse and heterogeneous nature of bidding and listing regulation data demands robust file parsing capabilities for data standardization. Unpredictable update frequencies necessitate support for external systems to push updates via API, enabling incremental parsing and chunking to ensure information timeliness. Documents contain extensive policy clauses and legal terminology. Their long-text characteristics challenge chunking strategies; overly short chunks may lose context, while overly long ones reduce retrieval precision. Specific identifiers like "policy number" require exact matching. Descriptive fields such as "scope" and "evaluation criteria" rely on semantic similarity retrieval. Furthermore, policy details vary significantly by province, requiring effective differentiation through tags or metadata to support precise retrieval based on region or product category.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk Length | 500–800 characters | Policy clauses often have strong contextual relevance. This length helps preserve complete semantics. |
Chunk Overlap | 70–100 characters | Ensures semantic continuity between adjacent chunks, reducing information loss. |
Recall Count | Top 8–12 items | Bidding and listing questions often involve multiple clauses, requiring more recall results for comprehensive judgment. |
Similarity Threshold | 0.75–0.85 | Policy interpretation demands high precision. A higher threshold filters out irrelevant results. |
Rerank Count | Top 5 items | Improves the ranking position of the most relevant results based on a higher recall count. |
File Parsing Timeout | 600 seconds | Policy documents may contain numerous charts and complex layouts, requiring longer parsing times. |
Common Pitfalls
- The knowledge base retrieval results contain many irrelevant clauses. This indicates a low semantic association between recalled items and the query. The
Similarity Thresholdis set too low, failing to effectively filter noise. - After uploading a new policy document, user queries still return old content. This indicates retrieval results do not include the latest information. The knowledge base is not configured to automatically trigger parsing and index rebuilding after external systems update content via the
addKnowledgeAPI. - The knowledge base retrieval module fails to operate in a workflow with tool calls. This indicates successful tool calls, but the knowledge base is not queried. The workflow design did not correctly configure the trigger conditions or priority for the knowledge base retrieval module, leading to it being overridden by the tool call logic.
Verification Steps
- Upload a typical bidding and listing policy document (e.g., PDF with charts). Check file parsing logs to confirm no parsing failures or timeouts, and that chunked content is as expected.
- Test with bidding and listing questions for different provinces and drug categories. Verify that retrieval results accurately recall policy clauses corresponding to the region or category. Observe if
Recall CountandRerank Countare applied as configured. - Dynamically add a new policy document containing specific keywords via API. Immediately query for those keywords to verify if the knowledge base promptly indexes and recalls the new content.
- Simulate user queries, such as "What are the application conditions for XX province's XX drug bidding and listing?". Check if the returned knowledge snippets precisely cover the key information in the question and evaluate their contextual completeness.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.