Data Characteristics for This Category
Tender listing documents include product registration certificates, production licenses, product manuals, inspection reports, clinical evaluation reports, enterprise qualifications, and authorization documents. They also include tender documents, winning bid notifications, and listing catalogs published by provincial and municipal tendering authorities. Data sources are extensive, covering the National Medical Products Administration, provincial and municipal pharmaceutical procurement platforms, medical institution procurement platforms, and internal enterprise management systems. Update frequency varies based on policy adjustments, product registration changes, and tendering cycles, typically quarterly or annually. However, policy documents and bidding information may be released in real-time. Document formats are diverse, with PDF, Word, and Excel being primary forms. PDF documents often contain scanned images, lacking directly extractable structured information. Key fields such as product generic name, approval number, manufacturer, dosage form/specification, winning bid price, listed province, and effective date often contain non-standard abbreviations, synonyms, and inconsistent units.
Constraints Imposed by These Characteristics on Multi-turn Conversations and Prompts
The unstructured, multi-source, and heterogeneous nature of tender listing documents requires multi-turn conversation systems to have robust document parsing capabilities, especially for OCR recognition and information extraction from scanned PDFs. Uncertain update frequencies necessitate knowledge bases that support incremental updates and version management to ensure the timeliness of referenced information. Inconsistent fields and units directly impact prompt accuracy, requiring standardization and entity linking during preprocessing to prevent model ambiguity in understanding user queries. For example, if a user asks for "the winning bid price of a certain product in Guangdong," the system must identify the administrative division corresponding to "Guangdong" and extract the correct "winning bid price" field from different document formats. Additionally, since documents involve regulations, technical details, and commercially sensitive information, conversations must strictly control information boundaries to prevent the model from generating inaccurate or non-compliant answers.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
maxContext | 8000 tokens | Registration and declaration documents are often lengthy, requiring a larger context window to hold complete information. |
Chunk size (Segment Length) | 500 characters | Balances information completeness and recall efficiency, preventing long paragraphs from diluting key information. |
Recall count (Recall Count) | 8 items | Ensures coverage of multiple relevant document segments, improving the comprehensiveness of answers. |
Similarity threshold (Similarity Threshold) | 0.78 | Filters out low-relevance documents, reducing noise and improving answer accuracy. |
Rerank result count (Rerank Return Count) | 4 items | Selects the most relevant entries from recall results, optimizing the final context presented to the model. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Parsing large and complex documents, such as scanned PDFs, can take a long time. |
Three Common Pitfalls
- After entering and running a prompt, the system returns a
400 status code. This may be due to the prompt template referencing a non-existent knowledge base tag or variable name. - The AI's answer contains critical information, such as prices or approval numbers, that does not match reality. This is because the knowledge base data was not updated in time, or the document parsing failed to extract the latest data correctly.
- During a conversation, the model cannot accurately answer questions about a specific product's listing status in a particular province. This may be due to a lack of indexing for regional tender documents in the knowledge base, leading to a failure in recalling relevant information.
How to Verify Configuration
- Test whether the model can accurately reference core fields like product generic name, approval number, and manufacturer from the knowledge base for a batch of typical registration and declaration questions.
- Simulate user queries about the winning bid prices of specific products in different provinces. Verify whether the model can extract and present correct price information from multiple sources and check its timeliness.
- Check if the model can distinguish between different versions of product manuals or inspection reports in its answers and reference the corresponding version's information to verify the knowledge base's version management function.
- Randomly select some PDF documents containing scanned images for questioning tests. Confirm whether the OCR recognition and information extraction accuracy meet expectations and check if key fields are correctly identified.
The values provided are common starting points and should be measured against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.