Data Characteristics for This Category
Data for culture media and consumables primarily originates from vendor product manuals, technical specifications, batch test reports, and internal laboratory validation records. This data updates infrequently, typically with product iterations or batch changes, approximately quarterly or semi-annually. Document structures are predominantly PDF, Excel, or Word, mixing unstructured text with tables. Key fields include product name, catalog number, lot number, manufacturing date, expiration date, storage conditions, main components, quality control indicators (e.g., osmolality, pH value, endotoxin levels), applicable cell types, recommended usage concentration, and specific certification information. Units often express component concentration in g/L or mg/L, pH values as units, and osmolality in mOsm/kg.
Constraints Imposed by These Characteristics on Tool Calling and Plugins
The low update frequency of culture media and consumables data means frequent data synchronization is unnecessary. The complexity of document formats requires tool calling to handle diverse, heterogeneous data sources, especially extracting table and text information from PDFs and scanned documents. Standardizing key fields like components and quality control indicators, and unifying units, presents a challenge for tool calling, requiring data cleaning and transformation. For example, different vendors may use varying names for the same component or inconsistent unit expressions, affecting subsequent accurate matching and pre-screening. Tool calling must parse these specific numerical values and perform comparisons, such as determining if a culture medium's pH value falls within a specific range or if endotoxin levels are below a threshold.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale for This Value |
|---|---|---|
tool_retrieval_mode | hybrid | Balances keyword matching and semantic similarity, improving recall for complex queries. |
max_tool_execution_time_seconds | 60 seconds | Most data extraction and cleaning operations complete quickly, avoiding excessive wait times. |
max_input_tokens_for_tool | 2048 tokens | Sufficient to cover key information summaries from most product manuals or batch reports. |
chunk_size | 800 characters | Balances contextual completeness with search efficiency, reducing redundant information. |
overlap_size | 100 characters | Ensures sufficient contextual linkage between segments, preventing critical information from being cut off. |
similarity_threshold | 0.75 | Balances recall and accuracy, reducing interference from irrelevant results. |
Three Common Pitfalls
- Tool call returns empty or incomplete results: This occurs when data extraction plugins lack sufficient parsing capability for unstructured documents, failing to correctly identify key fields or tables.
- Pre-screening results show unit errors or abnormal numerical comparisons: This happens when the data cleaning stage fails to standardize units from different sources, leading to incorrect numerical comparison logic.
- Historical queries cannot reuse previous tool call results: This is due to improper
tool_cache_strategyconfiguration, either disabled or set with too short a cache validity period.
How to Verify Correct Configuration
- Simulate various query scenarios using typical culture media and consumables product manuals to verify the tool's ability to accurately extract key data such as
main components,pHvalue, andendotoxin levels. - For specific pre-screening rules (e.g., requiring
pHvalues between7.0-7.4andendotoxin levelsbelow0.25 EU/ml), verify that the results returned after tool calling meet the expected screening conditions. - Check tool execution logs to confirm that
tool_execution_time_secondsdoes not exceed the configured threshold and that no error messages likeparsing_errororunit_conversion_failureappear.
The values provided are common starting points. Measure them against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.