Data Characteristics
Preclinical safety assessment data originates from experiment protocols, raw records, analysis reports, and GLP (Good Laboratory Practice) compliance statements. These documents are typically in PDF, Word, or scanned image formats. Data update frequency is relatively low, primarily occurring during project initiation, interim report submission, and final report archiving. Document structure is highly standardized, adhering to regulatory guidelines (e.g., FDA, NMPA). Documents include fixed fields such as study number, test article information, animal information, dosage, administration route, observation indicators, statistical data, and safety conclusions. Units use the International System of Units (e.g., mg/kg, IU/mL, kPa) and require high precision.
Constraints Imposed by Data Characteristics on Knowledge Base Retrieval and Recall
The highly structured nature and fixed fields of preclinical safety assessment documents require chunking to prioritize semantic completeness. This avoids separating critical data from descriptive text. The low update frequency means the knowledge base index reconstruction does not need to be frequent. However, each update must ensure data completeness and consistency. The large number of specialized terms and abbreviations requires pre-trained models to understand domain-specific vocabulary to improve recall accuracy. Strict requirements for data precision and units mean retrieval results must focus on extracting and displaying numerical information, ensuring correct unit matching, and preventing misinterpretations due to unit ambiguity.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size | 500–800 characters | Ensures a single segment contains a complete experimental design or result description, reducing semantic breaks. |
Overlap Length | 50–100 characters | Prevents important information from being split across different paragraphs, maintaining contextual continuity. |
Recall count | Top 5–8 entries | Preclinical safety assessment documents have high information density. Fewer retrieved items can cover most relevant content, reducing redundancy. |
Similarity threshold | Calibrate by actual measurement | Requires experimental determination based on the specific dataset to ensure highly relevant paragraphs are recalled while filtering out low-relevance noise. |
PARSE_FILE_TIMEOUT_SECONDS | 300 seconds | Large safety assessment reports may contain numerous charts and complex layouts, potentially requiring longer parsing times. |
maxContext | 4096 tokens | Ensures the model can handle longer contexts, understanding complex experimental designs and multi-indicator correlations. |
Common Pitfalls
- After knowledge base ingestion, unnecessary spaces appear between numbers and units in retrieval results. This happens because the parser introduces extra characters when processing specific layouts, affecting the precise matching of numerical values.
- Uploading Excel spreadsheet files fails. The current knowledge base primarily supports document formats like PDF and Word, lacking direct parsing capabilities for complex tabular data structures.
- Retrieval results include paragraphs irrelevant to the query. This occurs when the similarity threshold is set too low, leading to the recall of many low-relevance general descriptions.
How to Verify Configuration
- Query core safety assessment indicators and test article information using multiple different phrasings. Verify that recall results include all relevant and accurate data segments.
- Select several representative safety assessment reports, upload them to the knowledge base, and examine the parsed segment previews. Confirm the completeness of key data points and paragraphs.
- Simulate real-world problems. Query complex questions containing numbers and units. Validate the accuracy of numerical and unit matching in the returned results. Check for abnormal spaces.
The values provided are common starting points. Measure them against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.