Data Characteristics for This Category
Tender listing data primarily originates from various levels of drug and medical consumable centralized procurement platforms. This data is typically published as structured or semi-structured electronic documents, such as XML, JSON, PDF, or Excel spreadsheets. Update frequency depends on policy releases and procurement cycles, usually quarterly or annually. However, urgent procurements or special items may be adjusted at any time. Document content includes key fields like product name, manufacturer, registration certificate number, specifications, listed price, purchasing unit, and effective date. Price fields may involve differentiated information across provinces and medical institutions, while units cover various forms like minimum dosage unit and packaging unit.
Constraints Imposed by These Characteristics on "Reference and Traceability"
The diverse sources and frequent updates of tender listing data require the knowledge base to have efficient data ingestion and update mechanisms to ensure the timeliness of referenced information. Documents contain a large amount of structured data, necessitating precise field extraction capabilities to prevent information misalignment due to format differences. Overlapping or conflicting listing information from different provinces and batches challenges deduplication and prioritization of referenced content. The complexity of prices and units requires accurate contextual association during referencing to avoid ambiguity. Furthermore, the need to query historical listing records requires the knowledge base to support version management and timestamp traceability to handle policy changes and data iterations.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Segment Length) | 800–1200 characters | Ensures each knowledge block contains sufficient context, prevents truncation of critical information, and controls recall granularity. |
Chunk Overlap Length (Segment Overlap Length) | 100 characters | Maintains relevance between knowledge blocks and handles semantic dependencies across segments. |
Recall count (Recall Count) | 5–8 items | Balances recall precision and model processing load, covering different dimensions of listing information. |
Similarity threshold (Similarity Threshold) | 0.75–0.82 | Reduces false positives, focusing on tender listing information highly relevant to the query. |
Rerank result count (Rerank Return Count) | Top 3 items | Further refines recall results, providing the most relevant few pieces of information to the large model. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Accommodates parsing time for large PDF or Excel documents, preventing file processing failures due to timeouts. |
Common Pitfalls
- The model response does not reference any knowledge base content. This manifests as the absence of citation markers or an empty citation list in the response. Possible causes include a
Similarity threshold(Similarity Threshold) set too high, preventing recalled knowledge blocks from passing the threshold, or aRecall count(Recall Count) that is too low, failing to cover valid information. - Knowledge base references do not align with the actual response. This manifests as a logical disconnect between the response content and the cited sources. Possible causes include an inappropriate
Chunk size(Segment Length), leading to incomplete semantic knowledge blocks, or aRerank result count(Rerank Return Count) that is too low, preventing the model from obtaining sufficient diverse information for synthesis. - External systems cannot retrieve reference content. This manifests as an empty
referencesfield in the JSON data returned by the publishing channel. Possible causes include the knowledge base reference function not being enabled in the publishing channel configuration, or themaxContextparameter at the application layer limiting the length of reference information output by the model.
Validation Steps
- Use the "Test" feature in the FastGPT backend. Input typical tender listing-related questions and observe if the recalled knowledge blocks are accurate and complete. Verify their
Similarity Scorefalls within the expected range. - Check FastGPT's "Logs" or "Debug" interface. During each Q&A process, verify that the knowledge base's
Recall count(Recall Count) andRerank result count(Rerank Return Count) match the configuration and that no parsing or recall-relatederror codesappear. - Submit query requests through the actually deployed external interface. Check if the
referencesfield in the returned JSON structure contains valid reference information. Validate that thesourceandcontentfields of the references point to the correct knowledge base entries. - For specific tender listing data with frequent updates, perform regular queries to verify that the information in the knowledge base has been updated promptly and can be traced back to the latest version of the original document through references.
Note: The values provided are common starting points and should be measured against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.