Data Characteristics
Pharmacovigilance data during the lead optimization phase originates from preclinical study reports, early clinical trial (e.g., Phase I, Phase II) safety data, in vitro and in vivo pharmacology and toxicology research, and literature on similar compounds. Data updates frequently as projects progress, especially after new synthesis batches or experimental results. Document formats vary, including structured experimental reports (e.g., Excel spreadsheets, CSV files), unstructured researcher notes, PDF literature reviews, and exported toxicology database files. Key fields include compound identifiers (e.g., SMILES, InChIKey), dosage (units mg/kg or µM), observed indicators (e.g., ALT or AST elevation), adverse reaction descriptions (e.g., hepatotoxicity, nephrotoxicity), mechanism of action (e.g., CYP450 inhibition), and study time points (units days or hours).
Constraints on Knowledge Base Retrieval and Recall
The data characteristics of the lead optimization phase impose several constraints on knowledge base retrieval and recall. First, diverse document formats require robust multimodal processing capabilities. The system must extract precise numerical values from structured data and capture key adverse reaction descriptions from unstructured text. Second, compound identifiers enable retrieval based on chemical structure similarity. This requires integration or linking with external chemical information systems. Third, standardizing observed indicators and dosage units is crucial for retrieval accuracy. For example, mg/kg and µM require distinct and correct matching. Frequent data updates demand an efficient incremental indexing mechanism to avoid prolonged downtime. Furthermore, early research often involves many negative results or uncertain descriptions. This challenges the robustness of recall algorithms, which must handle incomplete or ambiguous queries.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Chunk Size) | 300–500 characters | Balances the completeness of a single adverse reaction description with retrieval granularity. Avoids noise from overly long contexts. |
Chunk overlap (Chunk Overlap) | 50 characters | Ensures critical information is not missed across segments, especially for sequential experimental data. |
Recall count (Recall Count) | 8–12 items | Limits the number of recalled results while ensuring coverage. Reduces complexity in subsequent processing, particularly when merging data from multiple sources. |
Similarity threshold (Similarity Threshold) | Calibrate based on measurements 0.75–0.85 | Balances recall and precision. Too low may introduce irrelevant toxicity data. Too high may miss potential risks. |
Rerank result count (Reranked Return Count) | 3–5 items | After reranking, focuses on a few most relevant results to the query, improving the quality of the final output. |
Embedding Model Version | text-embedding-ada-002 or newer | Ensures vectorization quality and improves the accuracy of semantic retrieval. Better understands specialized terminology. |
Common Pitfalls
- Query results are abnormally few, or even zero. A common cause is setting the
Similarity threshold(Similarity Threshold) too high. This strictly filters out relevant but slightly less similar document chunks. - Knowledge base retrieval takes too long, and dialogue response is slow. This usually occurs when the
Recall count(Recall Count) is set too high, or the knowledge base contains a large amount of data without effective index optimization. This overloads the retrieval backend. - Dialogue sometimes retrieves from the knowledge base and sometimes does not. This may be due to significant semantic differences between user input and knowledge base chunks. The
Embedding Modelmay fail to capture deeper associations, or an inappropriateChunk size(Chunk Size) may truncate key information.
Verification Steps
- For typical queries (e.g.,
hepatotoxicitydata for a specific compound), check if recalled results include all known relevant document chunks and verify their accuracy. - Use FastGPT's debugging interface to observe the
Recall count(Recall Count) andSimilarity Scorefor each knowledge base retrieval. Ensure they fall within the expected range. - Test the knowledge base's response speed with different query types (e.g., compound structure, adverse reaction description, mechanism of action). Evaluate performance under varying loads.
- After data updates in the knowledge base, check retrieval performance. Confirm that new data is retrieved promptly and accurately.
Note: The values provided are common starting points. Measure performance against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.