Data Characteristics
Data for preclinical safety evaluation (preclinical safety assessment) regulations primarily originates from official regulatory bodies like the National Medical Products Administration (NMPA) and the International Council for Harmonisation of Technical Requirements for Pharmaceuticals for Human Use (ICH). It also includes internal standard operating procedures (SOPs), internal guidelines, and experimental records from pharmaceutical companies. These documents are typically in PDF, Word, or scanned image formats. They feature a rigorous structure and contain extensive specialized terminology, dosage units (e.g., mg/kg, ppm), time units (e.g., days, weeks, months), animal strains, administration routes, and toxicity indicators (e.g., LD50, NOAEL). Document updates are relatively stable, usually occurring when regulations or guidelines are revised. Internal SOPs, however, may undergo minor revisions based on practical operations and accumulated experience. Fields and units are highly standardized and specialized.
Constraints on Multiturn Conversation and Prompts
The specialized and rigorous nature of preclinical safety evaluation regulations imposes high demands on the accuracy of multiturn conversations and prompt construction. Numerical information like dosages and times in the documents requires the model to accurately understand and cite them in conversations, avoiding unit confusion or misinterpretation of values. Due to the typically large volume of regulatory and SOP texts, and numerous cross-references, multiturn conversations need robust context management capabilities to support users in navigating and tracing information across different clauses. Additionally, the periodic nature of regulatory updates means the knowledge base requires regular maintenance to ensure the model always answers based on the latest valid information. The dense presence of specialized terminology requires prompts to guide the model in maintaining terminological accuracy in its responses, avoiding generalized explanations. The model also needs some ability to understand and convert colloquial user queries.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Chunk Size) | 500-800 characters | Ensures each knowledge chunk contains sufficient context while avoiding excessive length that could lead to information overload and reduced retrieval efficiency. |
Recall count (Retrieval Count) | Top 8 | Regulations and SOPs often have many cross-references; increasing the retrieval count helps cover a broader range of relevant information. |
Similarity threshold (Similarity Threshold) | 0.75-0.82 | Ensures the retrieved content is specialized and relevant, excluding generalized or imprecise matches. |
Rerank result count (Reranked Return Count) | Top 3 | Further refines the most core items while maintaining relevance, improving answer precision. |
maxContext | 4096 tokens | Accommodates the complexity and specialized nature of regulatory texts, providing an ample context window for multiturn conversations. |
temperature | 0.1-0.3 | Reduces model divergence, ensuring answers are fact-based and adhere to the rigorous requirements of regulatory frameworks. |
Common Pitfalls
- Phenomenon: The model's answer confuses dosage units, for example, misinterpreting mg/kg as g/kg. Reason: The original document contains various unit representations. Knowledge chunking failed to effectively differentiate them, leading the model to inaccurately identify them during generation.
- Phenomenon: A user asks, "Which clause mentions animal welfare?" The model responds, "No permission to operate this conversation record." Reason: The knowledge base lacks explicit permission management or citation tagging for regulatory clauses. When the model attempts to cite, it triggers a generic permission error message.
- Phenomenon: In a multiturn conversation, subsequent user questions are disconnected from the previous turn's context, and the model's responses lack coherence. Reason: The
maxContextparameter is set too low, preventing the model from effectively maintaining a longer conversation history, or the prompt design fails to effectively guide the model to focus on historical conversations.
How to Verify Configuration
- For a series of complex questions involving specialized terminology, dosages, and time units, check if the model's answers accurately cite values and units. Compare them with the original documents to verify accuracy.
- Simulate multiturn conversation scenarios. Test whether the model maintains contextual coherence and accurately answers subsequent questions when users trace and navigate between different regulatory clauses.
- Develop test cases to cover knowledge base updates after regulatory changes. Verify if the model can answer based on the latest valid information and provide correct explanations for outdated or superseded clauses.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.