Data Characteristics
R&D documents for culture media and consumables originate from internal experimental reports, supplier technical specifications, industry standard documents, and public literature. Updates are relatively stable, typically occurring with product batch changes, new material introductions, or standard revisions. Documents are primarily semi-structured, commonly in PDF, Word, and Excel formats. They contain extensive experimental data, formulation details, performance parameters, and quality control standards. Field characteristics include, but are not limited to: component name, CAS number, concentration (units: g/L, %, mM), batch number, storage conditions (unit: ℃), expiration date, purity, pH value, osmolarity (unit: mOsm/kg). Complex chemical structures and biological assay indicators are also present.
Constraints Imposed by These Characteristics on Multi-Turn Conversations and Prompts
The semi-structured nature of culture media and consumable documents means users may ask precise numerical queries (e.g., "What is the concentration of a certain component?") or vague descriptive questions (e.g., "Which culture media are suitable for cell culture?") in multi-turn conversations. This requires the model to accurately identify entities and numerical values and handle unit conversions. The document update frequency necessitates regular knowledge base maintenance and incremental updates to ensure the timeliness of conversation answers. Complex chemical formulas and biological indicators place higher demands on FastGPT's Embedding model selection and Chunk splitting strategy, requiring that these critical pieces of information do not lose semantic meaning during splitting. Simultaneously, users may frequently switch query targets in multi-turn conversations, requiring the system to accurately understand context and avoid "AI conversation getting stuck" or "unable to chat normally" scenarios.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk Size | 500–800 characters | Balances semantic completeness and recall efficiency, preventing long texts from diluting key information. |
Recall Count | Top 8–12 | Ensures rich context in multi-turn conversations, covering potentially relevant information. |
Similarity Threshold | 0.75–0.85 | Balances recall accuracy and false negatives, filtering out irrelevant document snippets. |
Rerank Count | Top 5 | Focuses on presenting the most relevant core information, improving conversation quality. |
maxContext | 4096 tokens | Accommodates long conversation scenarios, ensuring effective memory and understanding of historical conversations. |
Embedding Model | text-embedding-ada-002 or higher | Enhances understanding of specialized terms like chemical structures and biological indicators. |
Three Common Mistakes
- LaTeX format fails to render in conversations, displaying only raw code. This occurs if the integrated application frontend does not include a LaTeX rendering library, or if the
messagefield content is not processed for rendering. - During API calls, conversation logs show title content, but the actual return results lack critical information. This may be because the output of intermediate AI conversations in the
workflowis not correctly configured as the final workflow output. - The system remains stuck in an "AI conversation" state for an extended period, unable to chat normally. This may be due to insufficient
GPUresources in theFastGPTdeployment environment, or aLLMserver response timeout, leading to anHTTP 504error.
How to Verify Configuration
- Use the FastGPT debug preview feature to ask the model questions about culture media component concentrations or consumable batch numbers. Check the accuracy and completeness of the returned results.
- Simulate multi-turn conversations, progressively delving into a specific culture medium's performance indicators or storage conditions. Observe if the model maintains contextual consistency and provides relevant answers.
- Randomly select more than 10 R&D documents not used for training. Import them into the knowledge base and query their content to evaluate if the key information recalled by the model highly matches the document content.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.