Data Characteristics for This Category
Raw data for bidding and listing products originates from procurement platforms and public notices from medical institutions at various levels. This data typically exists as structured or semi-structured documents, such as PDF procurement catalogs, Excel spreadsheets of winning bids, and website announcements. Update frequency is cyclical. National or provincial-level updates occur every six months to a year. Supplemental procurement at city or hospital levels may happen monthly or quarterly. Document content includes key fields like product name, manufacturer, registration number, specifications, listed price, procurement quantity, and bid status. Price fields usually include currency units, and quantity fields include units of measurement. This information appears in diverse formats within documents and requires standardization.
Constraints on Model Integration and Configuration
The cyclical update rhythm and diverse document formats of bidding and listing data require continuous updates and flexible parsing capabilities for model integration. For example, a semi-annual provincial listing update means the model's knowledge base needs regular large-scale incremental or full synchronization. Different PDF documents have varying layouts and table structures. This requires optimizing the model's file parsing capabilities for complex documents to accurately extract core fields like product names and prices. Additionally, multiple files may describe different information for the same product (e.g., listed price and actual procurement price). Model configuration must consider how to integrate information and prioritize it to avoid conflicts or omissions. The diversity of field units (e.g., "yuan/box", "ten thousand yuan", "pieces", "units") also requires the model to understand and process this unit information for accurate numerical comparisons and calculations.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
maxContext | 3000 Tokens | Bidding and listing product descriptions are often long, containing multiple technical parameters and pricing information. A larger context window maintains completeness. |
Chunk size | 800–1200 characters | Ensures each segment contains complete key product information, such as name, specifications, price, and manufacturer, facilitating subsequent vector retrieval. |
Recall count | Top 5 entries | Product queries may involve multiple similar specifications or different batches. Increasing the number of retrieved items helps cover more comprehensive information. |
Similarity threshold | 0.75 | Bidding and listing product names and specifications contain many similar terms. A higher threshold allows for more precise matching of user intent. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Large PDF bidding announcement files take a long time to parse, requiring a longer parsing timeout. |
UPLOAD_FILE_MAX_SIZE | 100 MB | Bidding and listing files may contain many images and tables, resulting in large file sizes. Support for large file uploads is necessary. |
Common Mistakes
- Model response indicates "link disconnected" as the completion reason: This usually occurs due to an expired or misconfigured API key from the model service provider, preventing FastGPT from establishing an effective connection with the external model.
- Product price or quantity fields are empty in the model's reply: This may be because the file parser failed to correctly identify specific table structures or field units in PDF or Excel files, leading to data extraction failure.
- Using the same prompt yields a less effective FastGPT response compared to direct use on the native platform: This often happens when the knowledge base's segmentation strategy or retrieval parameters are set incorrectly, failing to provide the most relevant knowledge to the model, which prevents the model from responding with sufficient information.
How to Confirm Proper Configuration
- Upload typical bidding and listing PDF files. Check if the file parsing results accurately extract key fields like product name, specifications, and price, paying special attention to numerical values with units.
- Conduct consultation tests for multiple product names and specifications with significant differences. Verify if the model can accurately retrieve and integrate information from different documents.
- Simulate user queries about specific product prices, suppliers, and bid status. Validate the accuracy and completeness of the model's responses and compare them with the original documents to confirm information accuracy.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.