Data Characteristics
Lead compound screening data originates from high-throughput screening reports, activity assay data, structure-activity relationship (SAR) studies, toxicology pre-assessment reports, and relevant patent literature. This data updates infrequently, typically with experimental progress or project milestones. Document structures are complex, containing extensive unstructured experimental records, chemical structure images, charts, and semi-structured data tables. Fields include Compound ID, CAS number, molecular structure descriptions, in vitro activity data (e.g., IC50, EC50 values), ADMET property prediction results, and supplier information. Units vary, covering molar concentrations (nM, µM), percentage inhibition (%), and various toxicology metrics (e.g., LD50, mg/kg).
Constraints on Vector Models and Indexing
The complexity of lead compound screening data imposes specific requirements on vector models and indexing. First, the mix of unstructured text and structured data means that text embeddings alone are insufficient to capture all information. Consider multimodal embeddings or preprocess structured data for fusion with text. Second, infrequent data updates mean that index rebuilding does not need to be frequent, but each update may involve a large volume of data. The many chemical structures, charts, and other image information in documents require vector models with image recognition and description capabilities, or text extraction via OCR technology. Diverse units and numerical values challenge the embedding of numerical fields; models must understand and differentiate numerical magnitudes across different units, avoiding simple string matching.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
chunk_size | 800–1200 characters | Balances context completeness with vector model processing efficiency, preventing information fragmentation. |
overlap_size | 100–200 characters | Ensures contextual continuity and reduces information loss at segment boundaries. |
embedding_model | text-embedding-v2 or Doubao-embedding | Requires support for Chinese and English, with good comprehension of scientific literature. |
recall_top_k | 15–25 items | Increases the probability of recalling relevant document snippets, covering more potential information points. |
similarity_threshold | 0.78–0.85 | Balances recall and precision, filtering out irrelevant low-similarity results. |
metadata_fields | CompoundID, ActivityValue, TargetProtein | Combines structured metadata for filtering and refined retrieval, improving accuracy. |
Common Pitfalls
- Knowledge base progress halts after switching vector models, preventing a return to the original model. This typically results from a blocked backend task queue or failed model configuration validation. Check the API configuration and status of the model channel.
- When uploading PDF documents with many images and complex charts, some content is not indexed. This occurs because default text extraction capabilities are insufficient for processing image-based information. Enhance OCR or multimodal processing capabilities.
- Retrieval queries for specific compound activity values or ADMET properties fail to accurately recall documents containing precise numerical values. This may stem from ineffective embedding of numerical data or insufficient semantic understanding of number-unit combinations by the model.
Verification Steps
- Upload a batch of test documents containing typical lead compound data (e.g., IC50 values, molecular structure descriptions) and verify correct parsing and indexing in the knowledge base.
- Perform retrievals for specific Compound IDs, activity value ranges, or target names. Validate the relevance and accuracy of returned results, especially for queries involving structured information.
- Simulate common questions during regulatory submission preparation, such as "What is the IC50 of compound X for target Y?". Observe whether FastGPT recalls snippets containing precise numerical values and source documents from the index.
The values provided are common starting points. Measure them against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.