Data Characteristics for This Category
In the biomedical field, quality documents for culture media and consumables primarily include product specifications, COAs (Certificates of Analysis), SOPs (Standard Operating Procedures), and technical specifications. These documents are typically in PDF, Word, or scanned image formats. Data sources include official vendor materials, internal experimental records, and third-party testing reports. Update frequency is relatively low, occurring mainly when product batches change, suppliers are updated, or regulatory requirements are adjusted. Document structures usually contain fields such as product name, batch number, manufacturing date, expiration date, storage conditions, key ingredients, test indicators and results, quality standards, and usage instructions. Units involve concentration (e.g., g/L, %), pH value, temperature (℃), purity (%), and dimensions (mm). Unit inconsistencies may exist across different batches or suppliers.
Constraints Imposed by These Characteristics on Multi-Turn Conversations and Prompts
The characteristics of culture media and consumables quality documents impose specific requirements on multi-turn conversation and prompt design. A low document update frequency means manageable maintenance costs after knowledge base construction, but retaining and retrieving historical batch data is crucial. Diverse document formats, especially scanned images, require file parsing capabilities to effectively process text information from images and recognize tables and complex layouts. The coexistence of structured and semi-structured fields, along with unit variations, necessitates prompts that guide the model to perform precise data extraction and comparison, preventing result deviations due to unit confusion. For example, if a user queries "the pH value of a certain batch of culture media," the model must accurately identify the batch number and extract the corresponding pH value from the COA, while also confirming unit consistency. In multi-turn conversations, users may progressively refine query conditions, for instance, first asking "storage conditions for culture medium A" and then "expiration date for culture medium B." This requires the conversational system to maintain context and perform effective intent recognition and entity extraction.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
chunkSize | 800 characters | Balances document structural integrity with model processing capacity, preventing context loss. |
overlapSize | 100 characters | Ensures semantic continuity between segments, improving retrieval recall. |
maxContext | 16000 tokens | Addresses the need for long contexts in multi-turn conversations, handling complex queries. |
recallTopK | 10 items | Increases the breadth of initial retrieval, enhancing the likelihood of recalling relevant information. |
similarityThreshold | 0.75 | Filters for highly relevant document segments, reducing noise interference. |
rerankTopN | 5 items | Refines the final results, improving answer accuracy and user experience. |
Three Common Pitfalls
- After uploading a file in a conversation, the system displays
404 Invalid URL (POST /api/chat/completions). This occurs because theUPLOAD_FILE_MAX_SIZEparameter is set too low, causing file upload failure and preventing subsequent processing. - Knowledge base retrieval tests are normal, but content is occasionally not retrieved in conversations. This may be due to
similarityThresholdbeing set too high, filtering out document segments with slightly lower semantic similarity to the user's query, preventing them from entering the context. - The model confuses testing data from different batches in its response, for example, mixing the pH value of batch
20230101with the expiration date of batch20230201. This indicates insufficient information extraction and binding constraints for key entities like batch numbers in the prompt.
How to Verify Configuration
- Upload culture media and consumables documents in various formats (PDF, Word, scanned images) and check if the text content, especially tabular data, is correctly recognized in the knowledge base.
- Conduct multi-turn question-and-answer tests for specific batches and indicators. Verify if the model can maintain context and accurately extract information during a conversation, for example, asking "What is the protein content of batch X culture medium?" and then following up with "What is its storage temperature?"
- Check the model's response to numerical queries involving different units, for example, when asking "What is the pH range?", verify if the model can standardize or clearly label the units.
Note: The values provided are common starting points. Measure against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.