Data Characteristics
R&D documents from the lead optimization phase include compound synthesis reports, in vitro activity screening data, in vivo pharmacodynamics reports, ADMET (Absorption, Distribution, Metabolism, Excretion, Toxicology) evaluation reports, and Structure-Activity Relationship (SAR) analysis reports. Data sources are typically files exported from Laboratory Information Management Systems (LIMS), Electronic Lab Notebooks (ELN), and various specialized analysis software. Document updates are frequent, especially during SAR iterations, with new compound synthesis and test data continuously added.
Document structure typically includes:
- Compound number
- Molecular structure (SMILES or InChI)
- Experimental conditions
- Test results (e.g., IC50, EC50, Cmax, T1/2)
- Units (e.g., nM, µM, mg/kg, h)
- Metadata such as experimenter and date.
Reports often contain numerous charts, such as dose-response curves and pharmacokinetic curves. Key data from these charts usually appear in tabular format.
Constraints from "Context and Tokens"
The continuous update nature of lead optimization documents requires the knowledge base to efficiently handle incremental data and ensure context consistency between old and new information. Compound structures, experimental conditions, and quantitative data in documents are critical for precise matching and inference in a RAG (Retrieval-Augmented Generation) system.
Specialized terminology and symbols challenge the performance of tokenizers and embedding models, potentially leading to inaccurate token segmentation or low-quality vector representations. ADMET and pharmacodynamics reports are often lengthy, containing multiple indicators and complex experimental descriptions. This directly impacts the effective context length available from a single retrieval. If maxContext is set too small, the model may not acquire enough information for comprehensive analysis, leading to truncated answers or missing information.
The presence of multiple data types (text, tables, structural formulas) requires effective handling of heterogeneous information during context integration. This avoids dilution or neglect of critical data during tokenization.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Chunk Length) | 800–1200 characters (characters) | Balances the completeness of lengthy reports with the focus of individual information. Avoids excessive splitting of single compounds or experimental results. |
Chunk Overlap Length (Overlap Length) | 100–150 characters (characters) | Ensures contextual continuity across paragraphs, especially when describing experimental procedures and results. Prevents critical information from being fragmented. |
Recall count (Recall Count) | Top 5–8 entries (items) | Lead optimization problems often require multi-faceted information support. Increasing the recall count can improve critical information coverage. |
Similarity threshold (Similarity Threshold) | 0.75 | Filters out low-relevance document fragments, improving retrieval efficiency and answer quality, and reducing unnecessary token consumption. |
Rerank result count (Reranked Return Count) | Top 3–5 entries (items) | After screening by the reranking model, returns the most relevant few items. This refines the context and reduces model processing complexity. |
maxContext | 3000–4000 token | Ensures the model can accommodate multiple experimental report fragments. Supports complex SAR analysis and pharmacokinetic data comparison. |
UPLOAD_FILE_MAX_SIZE | 100 MB | Accommodates PDF or Word reports containing numerous charts and detailed data, ensuring seamless file uploads. |
Common Pitfalls
- Slow response times or high token consumption: This manifests as slow responses or high costs. The cause is
Recall count(Recall Count) ormaxContextbeing set too high, leading the model to process unnecessary redundant information. - Missing key numerical values or experimental conditions in model answers: This manifests as incomplete or truncated answers. The cause is
Chunk size(Chunk Length) being too small, leading to critical data being split across different segments. Alternatively,Similarity threshold(Similarity Threshold) is too high, filtering out relevant but slightly less similar fragments. - Reranker model startup failure or Token validation failure: This manifests as logs showing
token validation failedor the container failing to start. The cause is an incorrect or expiredrerankermodel ACCESS Token configuration.
Verification Steps
- Upload a typical lead optimization report (e.g., a complete SAR analysis report). Check the knowledge base chunk preview to ensure compound numbers, experimental results, and key parameters are within the same chunk.
- Ask complex questions about pharmacodynamic data for a specific compound. Observe whether the model's answer accurately cites different metrics and units from multiple reports, and verify the completeness of the answer.
- Simulate high-concurrency query scenarios. Monitor system resource utilization and average response time to ensure performance meets requirements under the configured
maxContextandRecall count(Recall Count).
The values provided are common starting points and should be measured against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.