Data Characteristics for This Category
Culture media and consumables registration data comes from various sources. These include supplier product manuals, Certificates of Analysis (COA), Material Safety Data Sheets (MSDS), internal quality control reports, batch inspection records, stability study reports, and regulatory standards (e.g., Chinese Pharmacopoeia, ISO standards). Data updates are relatively stable, with concentrated updates occurring when new products launch or regulations change. Documents are primarily in PDF format, containing extensive tabular data, chromatograms, experimental data, and technical parameters. Specific fields include "Batch Number," "Expiration Date," "Sterility," "Endotoxin Content," "Cytotoxicity," and "Osmolality." Units involve biological and chemical measurements like "U/mL," "EU/mL," and "mOsm/kg."
Constraints Imposed by These Characteristics on Knowledge Base Retrieval and Recall
The characteristics of culture media and consumables data impose several requirements on knowledge base retrieval and recall. First, documents contain tabular data and chromatograms. Simple text segmentation might lose context, requiring enhanced processing capabilities for non-textual information. Second, data update frequency is relatively stable, but updates often involve core parameter changes. This demands efficient incremental update mechanisms for the knowledge base. Regulatory standard citations make the precision and authority of retrieval results crucial, requiring avoidance of false positives. Additionally, the presence of specific fields and units necessitates that vector models understand the semantic associations of these specialized terms to improve recall accuracy. For example, a query for "endotoxin limit" should recall relevant content containing "EU/mL" units.
Configuration Guidelines
| Configuration Item | Suggested Value | Rationale for This Value |
|---|---|---|
Chunk size (Segment Length) | 500-800 characters (characters) | Balances paragraph completeness with retrieval granularity. This avoids excessively long segments that dilute the topic or overly short segments that lose critical information. |
Chunk overlap (Segment Overlap) | 50-100 characters (characters) | Maintains contextual continuity, ensuring complete recall of information across segments. |
Recall count (Number of Retrieved Items) | 8-12 entries (items) | Improves recall coverage, especially when handling complex queries and multi-faceted information needs, ensuring more relevant documents are included. |
Similarity threshold (Similarity Threshold) | 0.75-0.85 | Balances recall precision and recall rate. This reduces interference from irrelevant documents while ensuring high-relevance documents are hit. |
Rerank result count (Number of Reranked Items) | 5 entries (items) | Further filters the most relevant document segments using a reranking model after initial retrieval, improving the user experience. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds (seconds) | Accommodates the parsing time for large PDF files, preventing file upload failures due to timeouts. |
Three Common Mistakes
- When uploading large PDF files, the system reports
upstream connect error or disconnect/reset before head. This occurs because file parsing takes too long, exceeding the default gateway or proxy connection timeout limit. - Retrieval results contain many documents not directly related to the query. This happens when the
Similarity threshold(Similarity Threshold) is set too low, causing low-relevance documents to be recalled. - A query for "endotoxin content" fails to recall segments containing specific values and units (e.g., "less than 10 EU/mL"). This is often due to the vector model's insufficient semantic understanding of specialized units and numerical values, or
Chunk size(Segment Length) being set improperly, leading to critical information being truncated.
How to Confirm Correct Configuration
- Upload a culture medium COA file containing complex tables and specialized terminology. Check if file parsing is successful and verify the completeness of segmented content in the knowledge base.
- Execute a series of queries containing specialized keywords such as "Batch Number," "Expiration Date," and "Endotoxin Limit." Check if the retrieval results include highly relevant document segments and compare their consistency with the original document content.
- Perform a fuzzy query for a known regulatory requirement. Evaluate if the system can recall the corresponding regulatory provisions and check if the ranking of retrieval results meets expectations, with highly relevant documents appearing first.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.