Data Characteristics in this Category
Target discovery data originates primarily from biomedical literature, clinical trial reports, patent information, genomics and proteomics databases (e.g., UniProt, Gene Ontology), drug mechanism of action databases (e.g., ChEMBL, DrugBank), and internal experimental data. This data updates frequently, with new research and clinical data continuously released. Document structures typically include unstructured full-text research papers, structured database records, semi-structured experimental reports (e.g., PDF, DOCX), and supplementary materials containing tables and figures. Specific fields and units include numerous biomolecular identifiers (e.g., gene IDs, protein IDs), chemical structure information (SMILES, InChI), effect strengths (e.g., IC50, EC50, in nanomolar or micromolar units), and complex biological pathway descriptions.
Constraints Imposed by these Characteristics on FastGPT Deployment and Upgrade
The data characteristics in target discovery impose specific constraints on FastGPT deployment and upgrade. Data source diversity requires robust heterogeneous data integration capabilities, especially for the combined processing of unstructured text and structured databases. Continuous data streams mean the knowledge base needs to support incremental updates and version management to ensure information timeliness. Complex document structures and specialized fields, such as chemical structural formulas or biological pathway diagrams, may require customized parsers, as traditional text segmentation methods might not effectively extract key information. For example, models like text-embedding-ada-002 may require specific preprocessing steps or domain models when handling highly specialized data like biomolecular structures. Additionally, for numerical values with units, such as effect strengths, it is crucial to ensure correct identification and comparison during information extraction and retrieval to avoid misjudgments due to unit discrepancies.
Configuration Strategy
| Configuration Item | Suggested Value | Rationale |
|---|---|---|
UPLOAD_FILE_MAX_SIZE | 500 MB | Research papers and experimental reports can contain numerous charts and supplementary materials, leading to large file sizes. |
Chunk size | 800–1200 characters | Balances contextual completeness in biomedical text with embedding model processing efficiency, preventing key information truncation. |
Recall count | Top 10 entries | Target discovery often involves extensive related information; increasing recall helps cover a broader range of potential information. |
Similarity threshold | 0.75 | Sets a higher similarity threshold to ensure retrieval quality, given the precision requirements for biomedical terminology. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Parsing large PDF or DOCX documents can be time-consuming, requiring an extended timeout. |
Rerank result count | Top 5 entries | Refines the most relevant entries through reranking based on a high recall count, improving the accuracy of the final output. |
Three Common Mistakes
- Knowledge base related content does not appear in the answer, but the cited entries are correct: This may be due to the
maxContextparameter being set too small, preventing the model from fully utilizing all recalled contextual information when generating the answer. - After adding a custom plugin, input and output parameters are not displayed in the workflow: This usually occurs because the
inputsoroutputsfields in the plugin'splugin.jsonfile are not defined correctly, or there is a compatibility issue with FastGPT version (e.g.,v4.8.20-fix2) in parsing plugin metadata. - Knowledge base use reports an error "no available embedding model": This indicates that although the
text-embedding-ada-002model is configured in OneAPI,EMBEDDING_MODELin the FastGPT configuration file is not correctly pointed to or enabled, or the corresponding API Key has permission issues.
How to Confirm Correct Configuration
- Upload a PDF document containing biomolecular identifiers and effect strength data. Check if the knowledge base can correctly identify and extract these key information points after segmentation, especially
IC50values and their units. - Perform a target discovery search. Observe the distribution of
similarityscores in the search results and check if the cited content is highly consistent with key passages in the original document, ensuring the reasonableness ofRecall countandSimilarity threshold. - Call a custom plugin via the workflow. Verify that its input parameters correctly receive upstream data and that output parameters pass results as expected, confirming the validity of the plugin's
plugin.json.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.