High-Value Consumables Data Characteristics
High-value consumable data comes from various sources. These include manufacturer product manuals, technical whitepapers, compliance certifications (e.g., CE, FDA certificates), clinical application guidelines, and product catalogs. This data updates infrequently, typically a few times per year, coinciding with product iterations or regulatory changes. Document structures often involve PDF product manuals. Content includes product name, model, specifications, material, intended use, contraindications, usage instructions, sterilization methods, and storage conditions. Common fields include "Product Registration Number," "Batch Number," "Expiration Date," "Indications," "Packaging Unit," and "Minimum Sales Unit." Units include "mm," "g," "ml," or "pieces," "sets."
Constraints on Model Integration and Configuration
The low update frequency of high-value consumable data means real-time requirements for knowledge base construction are not high. However, accuracy and completeness are critical. Incorrect information can lead to severe medical risks. The prevalence of PDF documents requires robust PDF parsing capabilities for model integration, especially for tables and images. The specialized nature of fields and the precision of units challenge the model's ability to understand context and generate accurate answers. This requires more refined tokenization and entity recognition. For example, structured information like product registration numbers requires accurate extraction and association by the model. Additionally, significant variations in product specifications constrain the model's ability to handle long-tail queries and detailed comparisons.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk Length | 500–800 characters | Ensures each knowledge chunk contains sufficient product information without being too long, which could make it difficult for the model to focus. |
Recall Count | Top 5–8 items | Improves retrieval recall, covering various user query angles, especially with many product models. |
Similarity Threshold | 0.78–0.85 | Balances accuracy and recall, avoids irrelevant information interference, and matches approximate queries. |
Rerank Return Count | Top 3 items | Focuses on the most relevant product information, reduces model processing load, and improves response efficiency. |
maxContext | 3000–4000 tokens | Ensures the model can handle long queries with multiple product details and comparisons, particularly for product selection consultations. |
UPLOAD_FILE_MAX_SIZE | 100 MB | Accommodates potentially large PDF files for product manuals, preventing upload failures. |
Common Configuration Pitfalls
- Workflow model selection not taking effect, leading to unexpected output. This typically occurs because the model configured in the workflow node is not correctly bound or the selected model does not support the node's operation.
- Ollama-mounted local LLM model responses are disconnected from knowledge base content. This is due to mismatched OneAPI channel configuration parameters, preventing FastGPT from correctly calling or passing context.
- Rerank functionality not working or performing poorly. This usually happens when
rerank_api_keyorrerank_model_idare not correctly entered in system settings or knowledge base configurations.
Configuration Verification Steps
- Upload a typical product manual PDF. Check if the knowledge base chunk preview is accurate, especially if tables and key parameters are correctly extracted.
- Use queries containing keywords like product model, specifications, and indications. Verify if the model can recall relevant document snippets from the knowledge base.
- Ask about contraindications or storage conditions for a specific high-value consumable product. Cross-reference the model's answer with the product manual content.
- Simulate user product selection comparisons. Observe if the model can provide accurate comparison results based on different parameters (e.g., material, size).
The values provided are common starting points. Measure them against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.