Data Characteristics for This Category
Preclinical safety assessment data primarily originates from toxicology, pharmacokinetics, and pathology research reports. These are typically PDF-formatted experimental reports, study protocols, and SOP documents. These documents have complex structures, containing numerous charts, statistical data, specialized terminology, and abbreviations. Data update frequency is relatively low, usually occurring with project progress or batch completion, potentially once every few months. Fields include, but are not limited to, dose, administration route, animal species, observation indicators, test results, and statistical analysis results. Units involve milligrams per kilogram (mg/kg), milliliters per kilogram (mL/kg), days (day), hours (h), percentage (%), and require precise identification and processing. Some data also include unstructured expert opinions and conclusive descriptions.
Constraints Imposed by These Characteristics on "Deployment and Upgrade"
The large volume, complex format, and low update frequency of preclinical safety assessment data impose specific requirements on FastGPT's deployment and upgrade. First, the numerous charts and non-text content in large PDF reports necessitate high-performance document parsing capabilities to ensure complete information extraction. Second, the accuracy of specialized terminology and unit recognition directly impacts RAG retrieval quality, requiring targeted model fine-tuning or dictionary loading. The low update frequency means that one-time bulk import tasks are substantial, requiring ample computing resources and a stable network connection. Additionally, complex document structures may render default chunking strategies ineffective, requiring flexible adjustment of chunking parameters to avoid semantic discontinuity. During upgrades, data migration and index reconstruction time significantly increase, necessitating planned downtime windows and thorough verification of historical data compatibility.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
UPLOAD_FILE_MAX_SIZE | 500 MB | Preclinical safety assessment reports are often large, containing multi-page charts and detailed data. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Parsing large PDF documents is time-consuming; this prevents parsing failures due to timeouts. |
Chunk size | 800–1200 characters | Ensures the integrity of critical data and context in toxicology reports, preventing semantic fragmentation. |
Similarity threshold | 0.75–0.85 | Improves the precision of retrieval results and reduces interference from irrelevant or ambiguous information. |
Rerank result count | Top 5 entries | Preclinical safety assessment questions demand high accuracy, requiring a selection of a few high-quality retrieval results for reranking. |
maxContext | 4096 tokens | Accommodates the longer context dependencies in specialized reports, ensuring the model can understand complex logic. |
Three Common Mistakes
- Receiving a
413 Request Entity Too Largeerror when uploading large PDF files indicates that theUPLOAD_FILE_MAX_SIZEconfiguration is too small, causing the server to reject overly large files. - Model connection failures or abnormal responses after an upgrade occur because the API address, key, or model name in the
ONEAPIconfiguration does not match the new version requirements, or old version caching prevents connection information from refreshing. - Knowledge base query results contain excessive irrelevant information or lack critical data because the
Chunk size(chunk size) is set improperly, leading to incorrect semantic segmentation of documents, or theSimilarity threshold(similarity threshold) is too low, introducing noisy data.
How to Confirm Correct Configuration
- Upload a typical preclinical safety assessment report (e.g., a PDF with charts and toxicology data). Observe if the file parses and ingests correctly, and check if the ingested chunks are semantically complete.
- Ask questions about specific toxicology indicators or experimental conclusions within the report. Verify that the model accurately retrieves relevant passages and provides correct answers, ensuring the
similaritymetric of the retrieval results meets expectations. - Simulate a knowledge base update operation. Check if newly added or modified documents are processed correctly, and verify that the updated data can be retrieved and utilized in queries.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.