Data Characteristics in this Category
Data for lead compound screening primarily originates from drug target research reports, high-throughput screening results, in vitro pharmacodynamics data, toxicology assessment reports, and relevant patent literature. These documents are typically in formats such as PDF, Word, and Excel, with PDF being dominant. They often contain numerous tables, charts, and chemical structures. The data update frequency is relatively low, primarily occurring with new target discoveries, novel compound synthesis, or the release of evaluation results. Document structures are complex, frequently including sections like abstract, introduction, materials and methods, results, and discussion. Key fields include compound structure, IC50/EC50 values, solubility, bioavailability, target affinity, and ADMET properties. Units vary, including micromolar (µM), nanomolar (nM), and milligrams per kilogram (mg/kg), often accompanied by +/- error ranges.
Constraints Imposed by these Characteristics on "Document Parsing and Chunking"
Complex tables and embedded images in PDF documents challenge traditional text extraction tools, potentially leading to data loss or misalignment. Lengthy research reports and patent literature require parsers to effectively identify section structures to prevent context breaks. Chemical structures, as core information, are graphical in nature, making direct text parsing difficult and necessitating specialized image recognition or cheminformatics tools. The presence of units and error ranges requires chunks to maintain the association between these numerical values and their corresponding fields to avoid semantic fragmentation. Due to the infrequent data updates, knowledge base construction should focus on high-quality, one-time parsing to reduce future maintenance costs. The diversity of fields and complexity of units mean that a single, general chunking strategy is insufficient, requiring tailored chunking rules to ensure the completeness and retrievability of critical information.
Configuration Settings
| Configuration Item | Suggested Value | Rationale |
|---|---|---|
Chunk size (Chunk Length) | 800–1200 characters (characters) | Ensures each chunk contains sufficient context to cover complete experimental result descriptions or compound property lists. |
Chunk Overlap Length (Overlap Length) | 100–200 characters (characters) | Maintains contextual continuity, especially when tables, lists, or key conclusions span chunk boundaries, preventing information fragmentation. |
Enable Title Chunking | Enabled | Identifies document section and subsection titles, organizing related content together to maintain logical integrity. |
Parse Tables | Enabled | Lead compound data is often presented in tables. Enabling this extracts table data accurately, avoiding multi-row merging. |
Custom Separator (Custom Delimiter) | For Excel files, set to \n (newline character) | Excel source data often has one record per line. Using newline as a delimiter ensures each record is an independent chunk, preventing multiple records from merging into one. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds (seconds) | Provides sufficient parsing time for large PDFs or documents with complex structures (e.g., many images, tables), preventing parsing failures due to timeouts. |
Three Common Mistakes
- After document parsing, some table data appears with multiple rows merged into a single chunk, leading to information confusion. This occurs when table structures are not correctly identified or when an appropriate custom delimiter is not set for Excel source files.
- After uploading large PDF files, the system becomes unresponsive for an extended period or reports a parsing failure. This might be due to a
PARSE_FILE_TIMEOUT_SECONDSconfiguration that is too low, not providing enough processing time for complex documents. - Key compound structure information is missing from recall results, or it is separated from numerical data (e.g., IC50). This happens when image content is not effectively handled during chunking, or when the chunking strategy separates numerical values from their corresponding compound descriptions.
How to Verify Configuration
- After uploading typical documents (e.g., PDFs with complex tables, multi-row Excel files), check the number and content of chunks in the knowledge base. Ensure each chunk contains complete and logically coherent information, especially critical compound properties and experimental data.
- For different chunks in the knowledge base, use specific query terms (e.g., compound name, target name, IC50 value) to retrieve information. Verify the accuracy and completeness of the recall results, ensuring relevant information can be effectively retrieved.
- Through FastGPT's debugging interface, review document parsing logs to confirm no anomalies like
TimeoutorParsing Errorare present, especially for documents that previously failed to parse. - For documents with specific formats (e.g., custom delimiters), use the knowledge base's preview function to check chunk boundaries. This ensures custom delimiters are functioning as intended and prevents multiple records from merging into one chunk.
Note: The values provided are common starting points. They should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.