Data Characteristics
Lead optimization quality documents primarily include compound synthesis records, in vitro activity screening reports, in vivo pharmacokinetic (PK) data, preliminary toxicology assessments, and related Certificates of Analysis (CoA). Data sources are diverse, covering Laboratory Information Management Systems (LIMS), Electronic Lab Notebooks (ELN), and reports from third-party CROs. Update frequency is typically weekly or bi-weekly, dynamically updating with experimental progress. Document structure is semi-structured; for example, experiment reports contain fixed chapter titles, but the specific content is unstructured text. Fields involve compound structures (SMILES, InChI), batch numbers, concentration units (nM, μM), dosage (mg/kg), time points (h), activity inhibition rates (%), and quality standards (purity %, impurity limits %).
Constraints Imposed by These Characteristics on Workflow Orchestration
The multi-source and heterogeneous nature of lead optimization quality documents requires workflow orchestration to support various connectors and file format parsing during data ingestion. The semi-structured nature of documents necessitates customized preprocessing steps to accurately extract key field information from unstructured text, such as using regular expressions or custom parsing rules to identify compound batch numbers, experimental conditions, and results. The update frequency requires workflows to have scheduled trigger mechanisms to ensure the timeliness of knowledge base content. The specialized nature of fields and consistency of units are crucial for the accuracy of RAG retrieval results, requiring domain-specific fine-tuning of embedding models and unit validation of retrieval results. Additionally, documents like batch analysis certificates often contain tabular data, requiring workflows to support structured extraction of table content for precise question answering.
Configuration Guidelines
| Configuration Item | Suggested Value | Rationale |
|---|---|---|
maxContext | 4096 tokens | Balances context length and model processing efficiency, accommodating detailed experimental reports. |
Chunk size | 512 characters | Ensures individual segments contain complete experimental steps or result descriptions, minimizing information truncation. |
Recall count | Top 8 entries | Covers multiple relevant experimental records, improving the comprehensiveness of RAG recall. |
Similarity threshold | 0.75 | Balances recall precision and generalization ability, avoiding interference from irrelevant documents. |
Rerank result count | Top 3 entries | Focuses on the most relevant experimental data and conclusions, reducing user reading burden. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Accommodates parsing time for large experimental reports or PDF files containing complex charts. |
Common Pitfalls
- In the knowledge base search card, query results are empty after variable reference. This may be due to a mismatch between variable names and actual input parameters, or incorrect configuration of variable mapping in the knowledge base.
- After triggering a workflow in a session, interactive buttons fail to appear as expected. This is typically due to missing "send message" or "user interaction" node configurations in the workflow, or the message body format not meeting front-end rendering requirements.
- After calling an external service to generate a flowchart link, the workflow fails to display the image correctly. This may be because the link format returned by the external service is incorrect, or FastGPT's rendering component does not support that image link type.
Verification Steps
- Upload a typical experimental report containing compound structures and PK data. Observe whether the knowledge base correctly extracts and indexes all key fields, and check if field values and units are correct.
- Simulate user queries, such as "What is the in vitro activity of compound
XbatchYat concentrationZ?". Verify that RAG retrieval results include the correct experimental data and source document links. - Configure a scheduled workflow to simulate weekly data updates. Check if knowledge base content is automatically incrementally synchronized as planned, and verify the retrievability of newly added documents.
- Test the external API call functionality within the workflow. Check if API request parameters are correctly passed and if API responses are processed and returned as expected.
Note: The values provided are common starting points. Measure them against your own samples for optimal performance.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.