Data Characteristics
Lead compound screening data comes from high-throughput screening reports, activity verification reports, structural confirmation reports, and preliminary pharmacokinetic (ADME) data. This data is often structured or semi-structured, found in lab records, spreadsheets, project reports, and scientific literature. Document structures vary, including standardized experimental templates, free-text descriptions, and embedded charts. Field names and units are highly specialized, such as IC50 values (nM or µM), LogP values, molecular weight (Da), and spectral data (e.g., NMR, MS reports). Data update frequency depends on R&D progress, updating with experimental batches or phase summaries.
Constraints from "Document Parsing and Chunking"
The diversity of lead compound screening data poses multiple challenges for document parsing. First, free-text descriptions with specialized terminology and complex molecular structures require parsers to accurately identify and extract key information. Second, embedded charts (e.g., SAR tables, spectra) in Word or PDF documents need OCR or image recognition to convert them into processable text. Third, data from different sources and formats, such as high-throughput screening results in Excel and pharmacokinetic reports in Word, require a unified parsing strategy to ensure information completeness. Finally, numerical data with multiple units, such as IC50 in nM and µM, require normalization after parsing to prevent unit confusion and data misinterpretation.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
UPLOAD_FILE_MAX_SIZE | 500 MB | Handles large experimental report files containing numerous charts and high-resolution spectra. |
Chunk size (Chunk Length) | 800–1200 characters | Preserves contextual integrity while preventing information overload from excessively long chunks. |
Chunk Overlap Length (Overlap Length) | 100 characters | Ensures sufficient overlap of key information between adjacent chunks for semantic association. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Accommodates complex documents with many charts and tables, requiring longer parsing times. |
maxContext | 4000 token | Adapts to specialized terminology and long sentences in lead compound screening reports, ensuring deep understanding. |
Similarity threshold (Similarity Threshold) | 0.75 | Ensures retrieved results are highly relevant to the query intent, reducing inaccurate recalls. |
Common Mistakes
- When uploading large experimental report files, a "file processing timeout" error might occur. This can happen if the
PARSE_FILE_TIMEOUT_SECONDSparameter is set too low, insufficient for processing documents with complex charts and tables. - Knowledge base query results show an IC50 value for a compound that differs from the actual report. The value is correct, but the unit is wrong. This occurs because the document parsing did not normalize units for specialized fields like
IC50. - After importing Excel high-throughput screening data, some rows display incompletely, or multiple rows merge into one. This happens when the
Custom Separator(custom delimiter) is not configured correctly, preventing FastGPT from accurately identifying and splitting each record in the table.
Verification Steps
- Upload a lead compound screening report containing various data types (text, tables, charts). Check if the knowledge base accurately and completely extracts key fields like compound names, activity data, and structural information.
- Perform specific queries on documents already imported into the knowledge base, such as "What is the IC50 value of compound X?". Check if the returned numerical values and units match the original document. Verify the
Similarity threshold(similarity threshold) is appropriate. - Upload an Excel file containing a large amount of experimental data. Check if the knowledge base independently splits and stores each data record, ensuring no data merging or omission. This verifies the effectiveness of the
Custom Separator(custom delimiter).
The values provided are common starting points and should be measured against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.