Data Characteristics
Early compound screening data originates from drug discovery literature, experimental records, high-throughput screening results, patent documents, and internal databases. These documents typically contain chemical structures, physicochemical properties, biological activity data, target information, and toxicity predictions. Update frequency is relatively low, primarily occurring when new compounds are synthesized, activity testing batches are completed, or patents are published. Document structures are complex, including both semi-structured experimental reports (e.g., charts, data tables) and unstructured text descriptions. Fields and units vary. For instance, chemical structures are represented by SMILES or InChI codes. Activity data often uses IC50 or EC50 (in nM or µM) and Ki values. Physicochemical properties include LogP and molecular weight (in g/mol).
Constraints Imposed by Data Characteristics on Deployment and Upgrade
The complexity of early compound screening documents places specific demands on system deployment. The co-existence of unstructured and semi-structured data means the parser needs to extract key information from images, tables, and plain text. This may require additional OCR or table parsing modules. Document update frequency is low, but single updates can involve large data volumes. Therefore, balancing incremental and batch updates is crucial for knowledge base synchronization. Standardizing fields and units is another major challenge; different document sources may use varying names or units, requiring preprocessing or normalization during parsing. The unique nature of chemical structure information may necessitate integrating specific cheminformatics toolkits to ensure accurate structure recognition and retrieval. During upgrades, key considerations include compatibility of old and new data structures, stability of parsing logic, and smooth transition of chemical structure processing modules.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
UPLOAD_FILE_MAX_SIZE | 500 MB | Early compound screening documents often contain many images and tables, leading to large individual file sizes. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Parsing complex documents, especially with OCR and table extraction, can be time-consuming. |
Chunk size | 800–1200 characters | Balances contextual completeness and retrieval granularity, preventing segments from being too long or too short. |
Similarity threshold | Calibrated by measurement | Requires fine-tuning based on actual data for chemical concepts and biological activity descriptions to ensure relevance. |
recallWindowSize | 5 | Contextual information for early compounds is highly correlated, so increasing the recall window is appropriate. |
embeddingModel | bge-large-zh-v1.5 | Balances understanding of Chinese and specialized terminology, improving vectorization quality. |
Common Pitfalls
- After container startup, logs show
Error: unknown option --init. This usually indicates an incompatibility between theinitoption in the Docker Compose service configuration and the current Docker version, causing container startup failure. - Knowledge base retrieval results do not match expectations, failing to recall early compound information that was previously retrievable. This might be due to an improper knowledge base index rebuilding strategy after an upgrade, or changes in the tokenizer or embedding model leading to inconsistent vector representations.
- After modifying
config.json, the model list does not update. Checking/app/data/configinside the container reveals that the configuration file is not synchronized. This typically indicates a Docker volume mounting issue, preventing the container from reading the new configuration from the host.
Verification Steps
- Upload an early compound screening report containing complex tables and chemical structure diagrams. Verify that the parsing results accurately extract key physicochemical properties and activity data, and that chemical structures are correctly identified.
- Perform a retrieval for specific early compound names or targets. Verify that the system recalls all relevant experimental records and literature snippets, and compare the relevance and completeness of the retrieved results.
- After a system upgrade, execute a predefined set of test cases covering different document types and query patterns. Ensure that all functional modules (including file upload, parsing, vectorization, and retrieval) operate correctly and that performance meets expectations.
The values provided are common starting points and should be measured against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.