Data Characteristics
Tendering and listing data originates from drug and consumable centralized procurement platforms, medical insurance bureau public announcements, and industry associations. Data updates frequently; some provincial platforms update daily, while national data is typically aggregated weekly or monthly. Documents are primarily structured tables, such as Excel or CSV, containing fields like drug name, generic name, manufacturer, dosage form, specification, packaging, listed price, purchasing unit, bid status, and procurement cycle. Some data may be published as PDF announcements, requiring unstructured information extraction. Price fields are usually precise to two decimal places, with diverse units (e.g., yuan/box, yuan/piece, yuan/tablet) and potential unit conversion relationships.
Constraints on Deployment and Upgrade
High-frequency updates require FastGPT to automate data synchronization and index rebuilding for timely information. Large volumes of structured table data demand efficient parsing and field mapping mechanisms. PDF announcements require robust OCR and information extraction capabilities, directly impacting resource consumption during file upload and processing. Diverse price units and potential conversion relationships necessitate that FastGPT's knowledge base supports complex numerical processing during storage and retrieval. This also requires higher precision from text embedding models to prevent retrieval errors due to unit confusion. Furthermore, large data volumes impose demands on FastGPT's storage capacity and retrieval performance, requiring careful hardware resource planning.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
UPLOAD_FILE_MAX_SIZE | 500 MB | Tendering and listing PDF announcements and Excel files are often large. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Accommodates parsing time for large tables or image-heavy PDF files. |
Chunk size | 800–1200 characters | Ensures each segment contains complete listing entry information. |
Similarity threshold | 0.75 | Improves retrieval accuracy and reduces interference from irrelevant listing information. |
Rerank result count | 10 entries | Prioritizes displaying the most relevant listing information. |
embeddingModel | text-embedding-ada-002 or higher | Enhances understanding of specialized terminology and price units. |
Common Pitfalls
- Irrelevant listing information appears in knowledge base retrieval results. This happens when
Similarity thresholdis set too low, recalling many low-relevance document segments. - Uploading large tendering announcement PDF files results in long processing times without response, eventually timing out. This occurs when
PARSE_FILE_TIMEOUT_SECONDSis insufficient for complex document parsing. - After a system upgrade, the login interface still displays the old version number. This is due to uncleared browser cache or outdated reverse proxy caching policies.
Verification Steps
- Upload a typical tendering and listing Excel file containing various specifications and price units. Verify that the knowledge base correctly parses all fields and values.
- Upload a PDF tendering announcement larger than
100 MBvia the FastGPT interface. Confirm that the file processes correctly and completes indexing. - Simulate a version upgrade. Clear the browser cache and access the system to confirm the interface displays the latest version number.
- Perform a retrieval for a specific drug name and specification. Validate the
similarityscores of the returned results and adjustSimilarity thresholdbased on business requirements.
The values provided are common starting points. Measure against your own samples to determine the most suitable configuration.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.