Data Characteristics
Data for media and consumables primarily originate from supplier Product Technical Data Sheets (TDS), Certificates of Analysis (CoA/CoC), Safety Data Sheets (SDS), and internal Quality Control (QC) records. These documents are often in PDF format. Content includes detailed chemical compositions, biological activities, physical parameters (e.g., pH, osmolality), storage conditions, batch-specific data, and manufacturing dates. Data update frequency is relatively stable, typically occurring when product formulations change or new batch data is released. Document structures combine tables and paragraphs. Field names vary by supplier, but common units include g/L, mM, % (w/v), IU/mL, and CFU/mL.
Constraints on Reference Tracing and Source Attribution from these Characteristics
The diversity of media and consumables data imposes specific requirements on reference tracing and source attribution. Complex PDF document structures demand robust text extraction capabilities to accurately identify key parameters in tables and descriptive information in paragraphs. Frequent updates to batch reports require the knowledge base to have efficient data synchronization and version management mechanisms, ensuring referenced information always relies on the latest batch data. Inconsistent field names challenge the semantic matching of Retrieval-Augmented Generation (RAG) systems, potentially leading to low relevance in citations. Furthermore, since this data directly impacts experimental validity, tracing to specific page numbers and paragraphs in original documents is critical for information reliability.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
maxContext | 3000 tokens | Ensures key paragraphs and tabular data from product specifications are included, preventing information truncation. |
Chunk size (Chunk Size) | 500 characters | Balances semantic integrity and retrieval efficiency, accommodating longer descriptive texts in technical documents. |
Recall count (Recall Count) | Top 5 entries | Increases the retrieval scope to cover more potentially relevant documents, given the high accuracy requirements for clinical trial pre-screening. |
Similarity threshold (Similarity Threshold) | 0.75 | Balances recall and precision, filtering out results with low semantic relevance to improve source attribution quality. |
Rerank result count (Reranked Return Count) | 3 entries | Focuses on the most relevant and information-dense sources, reducing manual screening effort for engineers. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Provides sufficient time to process large PDF files, preventing data loss due to parsing timeouts. |
Common Mistakes
- Generated results include outdated batch information. This occurs when the knowledge base fails to synchronize the latest supplier Certificate of Analysis (CoA) reports, leading to retrieval of old data versions.
- Cited sources point to ambiguous document areas, such as only providing a file name without a page number or specific paragraph. This happens when document parsing does not retain sufficient metadata (e.g.,
page_number,section_id). - Answers confuse parameters for similar products from different suppliers. This is due to insufficient identification and isolation of suppliers or product models when the knowledge base processes multi-source data.
How to Verify Configuration
- Randomly select 10 queries about specific media batch parameters. Check if the generated answer cites the latest CoA report for that batch and can pinpoint the specific page number within the report.
- Ask a question about the storage conditions of a consumable. Confirm that the cited source is the TDS for that consumable and verify that the cited
storage_temperaturefield value matches the original text. - Query using names of similar products from different suppliers. Observe if the generated results correctly differentiate and cite their respective product technical data sheets without confusion.
- Simulate a query involving complex tabular data. Verify that FastGPT can accurately extract key numerical values (e.g.,
pH_value,osmolality) from the table and trace them back to the table's location.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.