ImageNet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever et al.
586
Citations
13
Influential Citations
Journal of Cheminformatics
Venue
2020
Year
The technological advances of the past century, marked by the computer revolution and the advent of high-throughput screening technologies in drug discovery, opened the path to the computational analysis and visualization of bioactive molecules. For this purpose, it became necessary to represent molecules in a syntax that would be readable by computers and understandable by scientists of various fields. A large number of chemical representations have been developed over the years, their numerosity being due to the fast development of computers and the complexity of producing a representation that encompasses all structural and chemical characteristics. We present here some of the most popular electronic molecular and macromolecular representations used in drug discovery, many of which are based on graph representations. Furthermore, we describe applications of these representations in AI-driven drug discovery. Our aim is to provide a brief guide on structural representations that are essential to the practice of AI in drug discovery. This review serves as a guide for researchers who have little experience with the handling of chemical representations and plan to work on applications at the interface of these fields.
This review addresses a critical bottleneck in AI-driven drug discovery: the need for standardized, computer-readable molecular representations. As the abstract notes, the complexity of capturing all structural and chemical characteristics in a single syntax has led to a proliferation of representations, creating confusion for newcomers. By cataloging the most popular representations—many of which are graph-based—the paper provides a clear entry point for researchers from AI backgrounds who lack cheminformatics expertise.
The paper's significance is underscored by its 586 citations, indicating it has become a foundational reference in the field. It bridges two communities: machine learning practitioners who need to understand chemical input formats, and cheminformaticians who want to apply modern AI techniques. This cross-disciplinary guidance is essential for progress in areas like virtual screening, de novo drug design, and property prediction.
As a review, the paper does not present new experimental results. Its value lies in the synthesis of existing knowledge: it organizes decades of cheminformatics research into a digestible format for AI practitioners. The paper's impact is measured by its citation count (586) and its role in enabling subsequent work—many papers citing it likely used its guidance to select representations for their own AI models.
This paper has democratized access to chemical representation knowledge for the AI community. By lowering the barrier to entry, it has accelerated the integration of machine learning into drug discovery pipelines. Its practical, non-technical tone makes it accessible to graduate students and industry researchers alike. The focus on graph representations also foreshadowed the rise of GNNs in molecular modeling, which have since become a standard tool. As AI continues to transform drug discovery, this guide remains a key reference for ensuring that representation choices are informed and appropriate.
Alex Krizhevsky, Ilya Sutskever et al.
Ashish Vaswani, Noam Shazeer et al.
Douglas M. Bates, Martin Mächler et al.
Diederik P. Kingma, Jimmy Ba