Data Characteristics
Bidding and procurement data originates from government procurement platforms, pharmaceutical centralized procurement platforms, and internal procurement announcements from medical institutions. This data updates frequently, often weekly or monthly. Document structures are primarily unstructured text, containing numerous tables and lists. Examples include detailed procurement requirements, technical parameters, and supplier qualification criteria. Field names often use abbreviations and industry-specific terminology, such as "CRO" (Contract Research Organization) and "IRB" (Institutional Review Board). Data commonly mixes numbers and text, for example, "Injection(100mg/vials)2000vials" (Injection (100mg/vial) 2000 vials) or "项目周期:12Units months" (Project duration: 12 months).
Constraints on Knowledge Base Retrieval and Recall
High update frequency requires the knowledge base to have efficient incremental update and indexing mechanisms to ensure timely retrieval results. Complex table structures and mixed text within documents challenge segmentation strategies. Standard segmentation may disrupt logical connections within tables, affecting recall accuracy. Industry-specific terminology and abbreviations demand strong semantic understanding; otherwise, keyword matching errors may occur. The mix of numbers and text complicates text cleaning and vector representation, potentially introducing unnecessary spaces or delimiters that hinder precise matching. Additionally, bidding information is often lengthy and contains much non-core information, increasing recall noise.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Segment Length) | 500–800 characters | Balances the length of bidding announcements with paragraph completeness, preventing truncation of key information. |
Overlap Length | 100–150 characters | Ensures contextual continuity across segment boundaries, improving recall for cross-segment queries. |
Recall count (Recall Count) | Top 8–12 items | Balances coverage while reducing the processing load on subsequent models. |
Similarity threshold (Similarity Threshold) | Calibrate via testing | Determined experimentally based on business needs for precision and recall balance. |
Rerank result count (Rerank Return Count) | Top 5 items | Focuses on the most relevant results, improving user experience and subsequent processing efficiency. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Accommodates the potentially long parsing time required for large bidding announcement documents. |
Common Pitfalls
- Retrieval results contain many irrelevant bidding announcements. This occurs when segmentation strategies fail to effectively identify and filter non-core content from documents.
- Critical bidding project information is not recalled. This manifests as empty results when querying with clear keywords. This may happen if the knowledge base did not correctly process tabular data during ingestion, leading to unindexed key fields within tables.
- Abnormal spaces appear between numbers and text, causing precise matching failures. For example, searching for "100mg/vials" fails to match "100 mg/vials". This is due to a lack of standardization for such mixed formats during text preprocessing.
Validation Steps
- Select a set of test questions with known answers. Perform retrieval and check if the recalled results include correct and complete bidding information. Evaluate the ranking priority.
- Upload bidding documents containing complex tables and specific terminology. Review the knowledge base segment preview in the FastGPT backend to confirm the completeness of table content and key fields.
- For queries mixing numbers and text, verify the matching of corresponding information in the recall results. Ensure that spacing issues between numbers and units have been properly handled.
Note: The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.