Data Characteristics for This Category
Laboratory service data primarily originates from various experimental reports, analysis results, operating procedures, instrument calibration records, and project summary documents. Data updates frequently, typically generated after each experiment or at the completion of project phases. Document structures vary, including structured experimental data tables, semi-structured report templates, and unstructured experimental logs. Fields often include sample ID, batch number, experimental conditions (temperature, pressure, concentration), measurement parameters (absorbance, chromatographic peak area, mass-to-charge ratio), units (nM, μg/mL, ℃, kPa), and result judgments (pass, fail, abnormal).
Constraints Imposed by These Characteristics on Knowledge Base Retrieval and Recall
High-frequency data updates require the knowledge base to quickly index new documents to avoid recalling outdated information. Diverse document structures necessitate flexible text segmentation strategies to accommodate row records in structured data tables, sections in semi-structured reports, and key descriptions in unstructured logs. The presence of specific fields and units requires effective handling of numerical ranges, unit conversions, and specialized terminology matching during retrieval. For example, when querying "above 37 degrees Celsius," the system must understand "degrees Celsius" and compare it with temperature values in documents. Additionally, for experimental reports containing numerous charts and images, the quality of text extraction directly impacts subsequent retrieval effectiveness, requiring attention to OCR accuracy.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale for Recommendation |
|---|---|---|
Chunk size (Segment Length) | 500–800 characters (characters) | A single experimental step or result description in an experimental report typically falls within this length, ensuring semantic completeness. |
Chunk Overlap Length (Segment Overlap Length) | 50–100 characters (characters) | Ensures contextual relevance across segments, preventing critical information from being split. |
Recall count (Recall Count) | 7–10 entries (items) | Given the complexity of experimental results, increasing the recall count improves coverage and reduces omissions. |
Similarity threshold (Similarity Threshold) | 0.75–0.85 | Laboratory terminology demands high precision; too low a threshold may introduce irrelevant results, while too high may lead to omissions. |
Rerank result count (Reranked Return Count) | 3–5 entries (items) | After reranking, refine the number to provide the most relevant core information, enhancing user experience. |
UPLOAD_FILE_MAX_SIZE | 200 MB | Experimental reports often contain numerous images and embedded objects, resulting in large file sizes, requiring an appropriate increase in the limit. |
Three Common Pitfalls
- Phenomenon: When uploading large experimental report files, the system reports
offset is out of boundsorrequest entity too large. Reason: The file size exceeds the default FastGPT or reverse proxy server (e.g., Nginx) configuration limit, leading to failed file chunk uploads or rejected requests. - Phenomenon: Knowledge base retrieval accuracy is low; returned information does not match the query intent. Reason: Improper document segmentation strategy causes critical experimental parameters or measurement results to be split, or an unreasonable similarity threshold recalls a large number of irrelevant document segments.
- Phenomenon: System response is slow, especially when handling complex queries or large knowledge bases. Reason: The deployed model scale (e.g., 70B) is large, and performance optimization configurations are insufficient, such as a lack of GPU resources or an excessively large
maxContextparameter setting, leading to prolonged inference times.
How to Confirm Proper Configuration
- Select a typical experimental report, upload it to the knowledge base, check if the file is successfully parsed and segmented, and verify semantic completeness of segmented content using the backend preview function.
- For specific experimental conditions or results, use query statements containing specialized terminology and numerical units to perform retrieval. Evaluate the relevance and accuracy of recall results and observe the
Recall count(Recall Count). - Simulate high-concurrency query scenarios and monitor system response times to ensure retrieval speed meets requirements under expected load. Also, check the actual effect of the
maxContextparameter. - Regularly use representative queries to evaluate the performance of the
Similarity threshold(Similarity Threshold) in recall results and adjust according to business needs.
Note: The values provided are common starting points. Measure them against your own samples to determine the optimal configuration for your specific use case.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.