Data Characteristics for This Category
Culture media and consumables data originates from diverse sources. These include product specifications, technical manuals, Material Safety Data Sheets (MSDS), batch analysis reports from suppliers, and internal lab quality control records. Data update frequencies vary; new products or batch updates can lead to high-frequency changes, while core parameters for mature products remain relatively stable. Document structures typically include product name, catalog number, specifications, lot number, production date, expiration date, storage conditions, main components, intended use, quality standards, and performance indicators. Common fields include concentration units (e.g., g/L, mM), pH value, osmolarity (mOsm/kg), endotoxin level (EU/mL), sterility test results, and cell growth support capability.
Constraints Imposed by Data Characteristics on "Deployment and Upgrade"
The diversity and update frequency of culture media and consumables data impose specific requirements on FastGPT deployment and upgrades. First, multi-source heterogeneous data necessitates flexible data ingestion and cleaning mechanisms to ensure effective parsing of different document formats. Second, data changes from batch updates and product iterations require the knowledge base to support incremental updates, avoiding full re-imports. Specialized terminology, chemical names, and specific units of measurement in documents influence the selection and parameter tuning of text embedding models, requiring accurate model understanding and differentiation. Furthermore, query needs for time-sensitive information like expiration dates and storage conditions require considering time-dimensional data processing during deployment and maintaining data consistency during upgrades.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
UPLOAD_FILE_MAX_SIZE | 500 MB | Technical manuals and batch reports are often large and require upload support. |
maxContext | 1500 characters | Ensures coverage of key information fragments from product specifications. |
Chunk size | 300 characters | Guarantees semantic completeness of single segments, preventing truncation of critical information. |
Recall count | Top 8 entries | Increases the probability of recalling key information from multiple relevant documents. |
Similarity threshold | 0.75 | Balances accuracy and recall, reducing irrelevant results. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Prevents parsing timeouts when processing large PDFs or scanned documents. |
Three Common Mistakes
- When uploading a large number of batch reports, the interface displays "File parsing failed." This happens because the
PARSE_FILE_TIMEOUT_SECONDSparameter is too low, preventing the complex document from being parsed within the allotted time. - After upgrading the FastGPT version, the quality of results from the original product inquiry function declines, with missing key components or storage conditions. This occurs because the embedding model's training data in the old version differed significantly from the new version, leading to semantic understanding deviations.
- During deployment, container startup errors like
Error response from daemon:typically result from port mapping conflicts or insufficient memory allocation in thedocker-compose.ymlfile, preventing the service from starting correctly.
Verifying Correct Configuration
- Upload a batch of typical product specifications and batch reports. Check that all documents parse successfully and display correct metadata fields.
- For different batches and specifications of culture media, input inquiries containing key parameters (e.g.,
pH 值,Endotoxin). Verify the accuracy and completeness of the returned results. - Simulate a product update scenario by importing new product technical manuals. Verify the knowledge base's incremental update functionality and check that updated query results reflect the latest information.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.