Data Characteristics
Data for lead compound screening in pharmacovigilance typically originates from high-throughput screening reports, in vitro toxicity test reports, structure-activity relationship (SAR) analysis documents, and preliminary in vivo pharmacokinetic (ADME) data. This data updates frequently, potentially weekly or monthly, depending on research and development progress. Document structures are primarily structured or semi-structured. Common fields include compound ID, molecular structure (SMILES or InChI codes), experimental conditions, target sites, effect values (e.g., IC50, EC50, Ki), toxicity indicators (e.g., LD50, cytotoxicity concentration), solubility, and permeability. Field units vary, involving nanomolar (nM), micromolar (µM), milligrams per kilogram (mg/kg), and percentages (%).
Constraints Imposed by These Characteristics on Vector Models and Indexing
High-frequency data updates require the vector index to support efficient incremental updates, avoiding resource consumption from full rebuilds. Compound ID and molecular structure are core identification information; their uniqueness and distinguishability must be maintained during vectorization. Numerical fields like effect values and toxicity indicators need accurate embedding into the vector space to support precise similarity search and anomaly detection. Diverse field units necessitate that the vector model handles different magnitudes of data or that data undergoes standardization during preprocessing. The presence of semi-structured documents implies a need for flexible text chunking strategies to capture contextual information while maintaining the precision of structured data.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Chunk Length) | 256 characters | Balances structured information completeness and contextual relevance |
Recall count (Recall Count) | 10 entries | Balances breadth during initial screening with subsequent re-ranking efficiency |
Similarity threshold (Similarity Threshold) | Calibrated by actual measurements | Adjust based on specific toxicity event detection rate and false positive rate |
Rerank result count (Re-ranking Return Count) | 5 entries | Accurately locates key information, avoiding interference from irrelevant results |
embeddingModel | bge-m3 | Supports multiple languages and modalities, adapting to diverse data types |
incrementalUpdate | true | Addresses high-frequency data updates, reducing full index overhead |
Common Pitfalls
- After creating a knowledge base, no models are available in the text understanding model dropdown. This indicates an incorrect
API KEYorENDPOINT URLin the model channel configuration, leading to model registration failure. - Uploading a large number of structured data files results in excessively long index build times or timeouts. This occurs when the
PARSE_FILE_TIMEOUT_SECONDSparameter is set too low, failing to accommodate the time required for large file processing. - Similar compound recall results are inaccurate, failing to identify lead compounds with potential toxicity. This suggests the
Similarity threshold(Similarity Threshold) is set too loosely, or the vector model does not effectively encode the structural and activity characteristics of compounds.
Verification Steps
- Upload lead compound data with known toxicity information. Retrieve similar compounds and check if the recall results include the expected high-similarity compounds.
- Monitor CPU and memory usage of the indexing service to ensure resource consumption remains within acceptable limits during incremental data updates.
- In the knowledge base configuration interface, verify that the configured models, such as
bge-m3, are successfully loaded in the text understanding model list.
The values provided are common starting points and should be measured against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.