Data Characteristics for This Category
Clinical trial pre-screening data for culture media and consumables primarily comes from supplier technical specifications, product batch reports, Certificates of Analysis (CoA/CoC), Safety Data Sheets (SDS), and relevant scientific literature. These documents are typically in PDF, Word, or structured database formats. Data update frequency depends on product iterations and batch production, usually quarterly or semi-annually. Some critical batch data may update at any time. Document structures vary: technical specifications focus on product performance indicators and composition ratios, while batch reports detail production dates, lot numbers, specific test results, and expiration dates. Fields include, but are not limited to, cell growth curves, osmolality, pH values, endotoxin levels, specific ion concentrations, and sterility test results. Units involve mOsm/kg, pH, EU/mL, ug/mL, and others, requiring high precision.
Constraints Imposed by These Characteristics on Deployment and Upgrade
Data source diversity requires FastGPT to be configured with robust file parsing capabilities during deployment, especially for extracting tables and unstructured text from PDF and Word documents. The uncertainty of update frequency, particularly for critical batch data that can update at any time, necessitates system support for incremental updates and rapid index rebuilding to avoid service interruptions from full rebuilds. The complex document structures, specialized fields, and high-precision units demand that the model accurately identify and understand this information for effective pre-screening. For example, the ability to recognize professional units like EU/mL and mOsm/kg directly impacts pre-screening accuracy. Deployment environments must also consider data security and compliance, especially for batch production data, ensuring data isolation and access control. When upgrading FastGPT versions, pay close attention to improvements in the model's ability to parse new document types and more refined fields, and verify compatibility.
Configuration Settings
| Configuration Item | Suggested Value | Rationale |
|---|---|---|
UPLOAD_FILE_MAX_SIZE | 500 MB | Technical specifications and batch reports can contain extensive charts, graphs, and detailed data, resulting in large file sizes. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Complex PDF document parsing, especially those with many tables and images, requires longer processing times. |
Chunk size (Segment Length) | 800–1200 characters | Technical descriptions for culture media and consumables are often information-dense; longer segments help maintain contextual integrity. |
Recall count (Recall Count) | Top 10 | Clinical trial pre-screening demands comprehensive information. Increasing the recall count captures more relevant details. |
Similarity threshold (Similarity Threshold) | 0.75 | Ensures recalled information is highly relevant to the query, filtering out vague matches and improving pre-screening accuracy. |
RERANK_RETURN_COUNT | 5 items | Reranks initial recall results to ensure the most relevant and critical batch or specification information is prioritized. |
Common Pitfalls
- After an update, the model cannot accept uploaded files for answering: The
UPLOAD_FILE_MAX_SIZEparameter might not be configured correctly, leading to new file upload limits that differ from the old version and causing file upload failures. - No image indexing option when creating a new knowledge base after local deployment: The
FASTGPT_IMAGE_INDEXING_ENABLEDenvironment variable is not set totrue, or related dependencies are not installed correctly, preventing image indexing functionality from activating. - No response after Docker deployment: The
PORTmapping configuration in thedocker-compose.ymlfile is incorrect, or the firewall has not opened the port FastGPT is listening on, preventing external requests from reaching the container.
How to Verify Configuration
- Upload a PDF technical specification containing tables and professional units. Check if table content and fields like
pHandEU/mLare correctly parsed in the knowledge base. - Simulate a query including batch numbers and production dates. Verify FastGPT can accurately recall relevant information from multiple batch reports and provide an answer.
- Check system logs to confirm the
PARSE_FILE_TIMEOUT_SECONDSparameter is effective and no file processing failures due to parsing timeouts are recorded. - Through the FastGPT administration interface, verify that the
RERANK_RETURN_COUNTparameter's reranked return count matches the actual number of results in the query.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.