Data Characteristics
Bidding and procurement data originates from government drug procurement platforms, medical device procurement platforms, and hospital procurement websites. This data updates frequently, typically weekly or monthly. Document formats vary, including PDF-formatted bidding announcements, procurement files, and attachment lists, alongside some structured Excel spreadsheets. Document structures are complex, containing both general terms and specific technical parameters for particular drugs or devices, registration certificate information, manufacturer qualifications, prices, and supply cycles. Field names often use abbreviations or non-standard expressions. Units include milligrams (mg), milliliters (ml), tablets, units, boxes, yuan, and percentages (%), with potential mixing or conversion relationships between different units.
Constraints on Vector Models and Indexing
The diversity of bidding and procurement documents requires vector models to have strong generalization capabilities. This allows processing different document types and layout variations, preventing information loss due to formatting issues. High update frequency necessitates efficient incremental indexing for the knowledge base. It must quickly ingest new bidding information and update relevant vectors. The density of technical parameters and qualification information in documents makes fine-grained text segmentation critical. Overly long segments can dilute key information, while overly short segments can break semantic integrity. Non-standard field names and mixed units demand higher standards for entity recognition and standardization during text preprocessing. Vector models must capture semantically similar but differently expressed entities. Additionally, numerical information like price and supply cycle requires special attention to contextual semantics during vectorization to ensure retrieval relevance.
Configuration Guidelines
| Configuration Item | Suggested Value | Rationale |
|---|---|---|
Chunk size (Segment Length) | 300–500 characters | Balances the density of technical parameters in bidding documents with contextual completeness. This prevents key information from being cut off or diluted. |
Chunk Overlap Length (Segment Overlap Length) | 50 characters | Ensures semantic continuity between adjacent segments, improving retrieval recall, especially when key information spans segments. |
Recall count (Recall Count) | Top 8–12 items | Given the complexity of bidding documents, increasing the recall count improves coverage. This works in conjunction with subsequent re-ranking. |
Similarity threshold (Similarity Threshold) | 0.75–0.85 | Balances recall accuracy and relevance. A value too low may introduce noise; a value too high may miss important information. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Bidding documents often contain many images or complex tables, requiring longer parsing times. This prevents parsing timeouts. |
Rerank result count (Re-ranked Return Count) | Top 3 items | Refines initial recall results using a re-ranking model, focusing on the most relevant results. |
Common Pitfalls
- After importing files into the knowledge base, some critical field content is empty: This occurs because the document parser fails to correctly identify non-standard table or list structures in bidding documents.
- After local deployment, the log shows
Failed to connect to embedding model: This usually indicates incorrectONEAPI_URLorONEAPI_KEYconfiguration, preventing FastGPT from connecting to the vector model service or causing authentication failure. - Retrieval results contain many irrelevant general terms: This happens when the text segmentation strategy is too coarse, failing to effectively distinguish core technical parameters from descriptive background content in the document.
Verification Steps
- Upload a typical bidding announcement PDF file. Check if the segmented content in the knowledge base is complete and semantically coherent, paying special attention to technical parameters and qualification description paragraphs.
- For queries regarding specific drugs or devices, verify that the returned results include accurate registration certificate numbers, manufacturer names, and key technical indicators. Also, check their source documents.
- Simulate the addition of a new batch of bids. Observe the query response speed and accuracy of new data after the knowledge base updates. This assesses the effectiveness of incremental indexing.
- Use FastGPT's backend debugging tools to view the
Recall count(Recall Count) andSimilarity threshold(Similarity Threshold) for each query. Adjust parameters based on actual business needs until retrieval results are satisfactory.
The values provided are common starting points and should be measured against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.