Data Characteristics in This Category
R&D documents during the lead optimization phase typically include experimental reports, molecular structure data, ADME (Absorption, Distribution, Metabolism, Excretion) reports, toxicology study records, and synthesis route designs. These documents are commonly in PDF, Word, or structured data formats (e.g., SDF, CSV). Update frequency varies from daily, weekly, to real-time, depending on experimental progress. Document structures are diverse; some contain free-text descriptions, while others include tables and graphs. Common fields include compound ID, CAS number, molecular formula, molecular weight, activity data (e.g., IC50, Ki), physicochemical properties (e.g., LogP, solubility), ADME parameters (e.g., plasma protein binding, metabolic stability), and toxicity indicators. Units involve molar concentrations (nM, μM), mass (mg, g), and time (h, min). Different laboratories or reports may use varying unit representations.
Constraints Imposed by These Characteristics on "Multi-turn Conversation and Prompting"
The mixed structure of lead optimization documents presents challenges for multi-turn conversations. Free-text sections require robust semantic understanding to extract key information, while structured data demands precise field recognition and unit conversion. For example, a user might ask, "What is the IC50 of compound XYZ?" The system needs to identify the compound ID and locate the corresponding activity data across different reports. The real-time updates of ADME and toxicology data necessitate rapid synchronization and indexing capabilities for the knowledge base, ensuring conversation results are based on the latest data. Furthermore, the inconsistent use of units requires intelligent conversion or clarification during conversations to avoid misunderstandings. The complexity of multi-turn conversations also arises when users follow up with questions like, "What about the solubility of this compound?" or "Which known molecular structures are similar to it?" This requires the system to maintain context and perform associative retrieval and information integration from multiple document dimensions.
Configuration Settings
| Configuration Item | Recommended Value | Rationale for This Value |
|---|---|---|
maxContext | 8 turns | Lead optimization conversations often involve multi-step reasoning and information backtracking. 8 turns cover most scenarios, maintaining conversational coherence. |
Chunk size (Segment Length) | 800–1200 characters | Experimental report paragraphs are often long. Shorter segments risk losing context, while excessively long ones increase LLM processing burden and noise. |
Recall count (Recall Count) | Top 10 | Ensures sufficient potentially relevant document snippets are recalled initially, especially when information is dispersed. |
Similarity threshold (Similarity Threshold) | 0.78 | An empirical value that balances recall rate and accuracy, avoiding interference from irrelevant snippets. |
Rerank result count (Reranked Return Count) | Top 5 | Further filters recalled results, focusing on the most relevant information to improve final response quality. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Processing large experimental reports or structured data files can be time-consuming; this prevents parsing interruptions. |
Three Common Mistakes
- Symptom: A user queries activity data for a specific compound, but the returned result shows the field as empty or "not found." Reason: Document parsing failed to correctly identify or extract activity data fields from different formats (e.g., tables, graph annotations), or filtering occurred due to unit mismatches.
- Symptom: In a multi-turn conversation, a user follows up with a question about compound toxicity information, but the system fails to associate it with the previous compound ID and asks the user to re-provide it. Reason: The
maxContextparameter is set too low, leading to context loss and failure to maintain the historical state of the conversation. - Symptom: The AI conversation variable in the workflow cannot obtain molecular structure optimization results from code execution output. Reason: The code execution output was not written to the workflow variable in the expected data structure or format, preventing the AI module from correctly parsing it.
How to Confirm Correct Configuration
- Select 5–10 typical lead optimization scenarios. Simulate multi-turn conversations from an initial question to three follow-up questions. Verify that each response is accurate and contextually coherent, and check that the data mentioned in the responses aligns with the original documents.
- Upload 3-5 R&D documents each containing different formats (PDF, Word, CSV) and units (nM, μM, mg/kg). Check that key fields (e.g., IC50, LogP) are completely extracted into the knowledge base after document parsing and that unit conversions are correct.
- Monitor
PARSE_FILE_TIMEOUT_SECONDSandmaxContext-related events in FastGPT logs. Confirm that no timeout or context truncation warnings occur when processing large files or complex multi-turn conversations.
Note: The values provided are common starting points. Measure them against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.