Data Characteristics
Lead optimization data originates from high-throughput screening reports, structure-activity relationship (SAR) analysis documents, molecular docking results, toxicology prediction reports, and in vitro/in vivo efficacy experimental records. These documents often contain extensive structured and semi-structured data. Examples include compound IDs, molecular formulas, activity values (e.g., IC50, EC50), ADMET property prediction parameters, target information, and experimental conditions. Documents are updated frequently, especially during compound library screening and structural modification iterations. Fields and units are highly specific. Activity values commonly use nM or μM, molecular weight uses Da, and toxicity indicators might involve LD50 or LC50. Some data appears in tables, while other data is embedded within experimental descriptions.
Constraints on Document Parsing and Chunking
Lead optimization documents contain a mix of structured and semi-structured data. This requires parsers to recognize plain text and effectively process embedded tables. High update frequency necessitates efficient incremental update capabilities to avoid re-parsing already processed data. Specific critical fields, such as compound IDs and activity values, require the parsing model to accurately identify and extract data with specific units or formats. For example, "10 nM" and "10 μM" have significant semantic differences and require precise distinction. Fields like ADMET properties might appear as abbreviations, requiring a synonym dictionary for identification. Additionally, molecular structures are typically images. Document parsing needs optical character recognition (OCR) support to extract relevant structural information or annotations from images.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
UPLOAD_FILE_MAX_SIZE | 1000 MB | Accommodates large high-throughput screening reports or documents with numerous images/tables. |
Chunk size (Chunk Length) | 800–1200 characters (characters) | Balances context completeness and recall efficiency, suitable for longer experimental description paragraphs in lead optimization documents. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds (seconds) | Accounts for the time required to process large PDFs or documents with complex tables, preventing parsing timeouts. |
maxContext | 32000 tokens | Ensures sufficient capacity for critical experimental data, methods, and conclusions, meeting RAG model context window requirements. |
Enable Table Parsing | Yes | Many critical data points in lead optimization documents are presented in tables and require effective parsing. |
OCR Recognition Mode | High Accuracy | Ensures accurate recognition and extraction of molecular structures, chart labels, and experimental data from images. |
Common Pitfalls
- Uploading large PDF files results in a "413 Request Entity Too Large" error. This occurs when the server's file size limit is lower than the actual file size. Adjust the
UPLOAD_FILE_MAX_SIZEparameter. - Activity values or toxicity data fields are empty in parsed documents. This happens when the parser fails to correctly identify numerical patterns with units or when specific regular expressions are not configured for matching.
- Multi-level directory structures in Feishu or external PDF documents are not fully parsed, leading to content being unretrievable. This likely occurs because the parser defaults to processing flat text and does not adapt to the semantic structure of multi-level headings and directories.
Verification Steps
- Upload a lead optimization report containing complex tables and molecular structure images. Check if table data is complete in the parsed results and if fields and values correspond accurately.
- Randomly select key numerical values (e.g., IC50, LD50) from the document. Use keyword search to verify if these values are accurately recalled and if their units are correctly identified.
- Upload a multi-level R&D document. Review the parsed chunks to ensure the original document's logical hierarchy is preserved and that directory and heading information is effective.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.