ImageNet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever et al.
4.5k
Citations
298
Influential Citations
Nucleic Acids Research
Venue
2011
Year
ChEMBL is an Open Data database containing binding, functional and ADMET information for a large number of drug-like bioactive compounds. These data are manually abstracted from the primary published literature on a regular basis, then further curated and standardized to maximize their quality and utility across a wide range of chemical biology and drug-discovery research problems. Currently, the database contains 5.4 million bioactivity measurements for more than 1 million compounds and 5200 protein targets. Access is available through a web-based interface, data downloads and web services at: https://www.ebi.ac.uk/chembldb.
ChEMBL addresses a critical bottleneck in drug discovery: the lack of large-scale, high-quality, and openly accessible bioactivity data. Prior to ChEMBL, most bioactivity data was scattered across proprietary databases or buried in unstructured literature, making it difficult for researchers to build predictive models or perform large-scale analyses. By providing a centralized, manually curated repository, ChEMBL democratizes access to drug-target interaction data, accelerating both academic research and industrial drug development.
The database's emphasis on open data and standardization is particularly significant for the machine learning community. High-quality training data is essential for developing accurate predictive models, and ChEMBL has become a standard benchmark for tasks such as bioactivity prediction, compound-target interaction modeling, and ADMET property estimation. Its regular updates and broad coverage of targets and compounds ensure its continued relevance.
The paper reports that ChEMBL contains 5.4 million bioactivity measurements for more than 1 million distinct compounds and 5200 protein targets. These data are manually abstracted from the primary literature and curated to maximize quality. The database is accessible at https://www.ebi.ac.uk/chembldb. No specific performance metrics or comparisons to other databases are provided in the abstract.
ChEMBL has had a transformative impact on computational drug discovery and cheminformatics. It has enabled the development of machine learning models for predicting drug-target interactions, compound bioactivity, and ADMET properties. The database is widely used in both academia and industry, serving as a benchmark for evaluating new algorithms and as a training resource for deep learning models. Its open data philosophy has inspired similar initiatives and fostered a culture of data sharing in the drug discovery community.
Alex Krizhevsky, Ilya Sutskever et al.
Ashish Vaswani, Noam Shazeer et al.
Douglas M. Bates, Martin Mächler et al.
Diederik P. Kingma, Jimmy Ba