Data Characteristics
Preclinical safety assessment data primarily originates from pharmacology and toxicology research reports, GLP laboratory experimental data and spectra, domestic and international regulatory guidelines, and non-clinical safety data for marketed drugs. This data updates infrequently, primarily changing with new drug development progress and regulatory revisions. Document formats vary, including structured experimental data tables, unstructured research report text, PDF spectra and raw data records, and Word or Excel format regulatory documents. Fields include dosage, administration route, animal species, observation indicators (e.g., body weight, pathological findings, blood biochemical indicators) and their units. Descriptions of dose-response relationships and statistical analysis results are common. Data rigor and traceability are core characteristics.
Constraints on Knowledge Base Retrieval and Recall
Preclinical safety assessment data has both highly structured and unstructured components. This requires the knowledge base to effectively handle both tabular data and long text reports. The low data update frequency means knowledge base construction must focus on historical data completeness and version management. Recall does not require high timeliness, but demands extremely high accuracy. Extensive specialized terminology, abbreviations, and numerical information like dosages and units challenge text segmentation and the semantic understanding capabilities of embedding models. This requires ensuring the integrity of critical information. The strict formatting and citation rules in GLP reports require retrieval results to precisely pinpoint original sources for traceability and verification. Furthermore, preclinical safety assessment data is often sensitive information, making knowledge base security and access control important considerations.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Segment Length) | 500–800 characters (characters) | Balances the complete semantics of long texts with retrieval efficiency, preventing key information from being cut. |
Chunk Overlap Length (Segment Overlap Length) | 50 characters (characters) | Ensures contextual continuity and handles specialized terms or critical descriptions spanning segments. |
Similarity threshold (Similarity Threshold) | 0.75–0.85 | Preclinical safety assessment data requires high precision; increasing the threshold reduces irrelevant recall. |
Recall count (Recall Count) | Top 8–12 entries (top 8–12 items) | Ensures sufficient contextual information, covering multiple potentially highly relevant experimental reports. |
Rerank model (Rerank Model) | BGE-M3 or E5-Mistral-7B | Improves the accuracy of relevance ranking for specialized biomedical vocabulary. |
Max Context Tokens | 8192 | Accommodates long paragraphs and detailed descriptions in safety assessment reports, providing ample context. |
Common Misconfigurations
- Retrieval results include many irrelevant experimental data or regulatory clauses. This occurs when the
Similarity threshold(Similarity Threshold) is set too low, failing to effectively filter noise. - AI responses show incorrect understanding of dosages or units. This stems from
Chunk size(Segment Length) being too short, truncating the context of numerical values and units. - System logs show a much lower than expected number of recalled documents. This may be due to the embedding model's insufficient understanding of specialized terminology, failing to correctly match relevant knowledge blocks.
How to Verify Configuration
- Select several typical preclinical safety assessment queries. Check if the recall results precisely point to relevant research reports, experimental data tables, or regulatory clauses.
- Compare with original documents. Verify if key information (dosage, animal species, observation indicators) in AI responses based on the knowledge base matches the original text.
- Adjust
Similarity threshold(Similarity Threshold) andRecall count(Recall Count). Observe changes in the relevance and coverage of retrieval results until desired effects are achieved. - Query specific specialized terms or abbreviations. Verify if the knowledge base accurately identifies and recalls documents containing these terms.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.