Journal Article
Machine Learning

Computational approaches streamlining drug discovery

Anastasiia Sadybekov(University of Southern California), Vsevolod Katritch(University of Southern California)
April 26, 2023Nature1,192 citations

1.2k

Citations

14

Influential Citations

Nature

Venue

2023

Year

Abstract

Computer-aided drug discovery has been around for decades, although the past few years have seen a tectonic shift towards embracing computational technologies in both academia and pharma. This shift is largely defined by the flood of data on ligand properties and binding to therapeutic targets and their 3D structures, abundant computing capacities and the advent of on-demand virtual libraries of drug-like small molecules in their billions. Taking full advantage of these resources requires fast computational methods for effective ligand screening. This includes structure-based virtual screening of gigascale chemical spaces, further facilitated by fast iterative screening approaches. Highly synergistic are developments in deep learning predictions of ligand properties and target activities in lieu of receptor structure. Here we review recent advances in ligand discovery technologies, their potential for reshaping the whole process of drug discovery and development, as well as the challenges they encounter. We also discuss how the rapid identification of highly diverse, potent, target-selective and drug-like ligands to protein targets can democratize the drug discovery process, presenting new opportunities for the cost-effective development of safer and more effective small-molecule treatments.

Analysis

Why This Paper Matters

This paper, published in Nature in 2023 with over 1,100 citations, marks a pivotal moment in computational drug discovery. It captures the tectonic shift from traditional experimental methods to AI-driven approaches, driven by the convergence of massive datasets (ligand properties, 3D structures), abundant computing power, and the emergence of on-demand virtual libraries containing billions of drug-like molecules. For AI practitioners, this review is essential reading because it maps the landscape where machine learning—particularly deep learning—is becoming a core enabler, not just an auxiliary tool.

The paper’s significance lies in its comprehensive synthesis of two complementary paradigms: structure-based virtual screening (leveraging 3D target structures) and deep learning predictions of ligand properties and activities (which can bypass the need for receptor structures). By highlighting how these methods can be synergistically combined, the authors provide a roadmap for researchers aiming to accelerate the identification of potent, selective, and drug-like compounds. This democratization of drug discovery could lower barriers for smaller labs and startups, potentially reshaping the pharmaceutical industry.

Technical Contributions

The paper’s main technical contributions are its systematic review and categorization of computational approaches:

  • Gigascale virtual screening: Methods that can efficiently screen billions of compounds against protein targets, using fast docking algorithms and iterative screening to reduce computational cost.
  • Deep learning for ligand properties: Neural networks that predict ADMET (absorption, distribution, metabolism, excretion, toxicity) properties, binding affinities, and target activities directly from molecular structures, often outperforming traditional QSAR models.
  • Synergistic integration: Combining structure-based and ligand-based methods to improve hit rates and reduce false positives, e.g., using deep learning to pre-filter libraries before docking.
  • On-demand virtual libraries: The use of enumerated or fragment-based combinatorial libraries (e.g., REAL, Enamine) that can be synthesized on demand, bridging the gap between virtual hits and experimental validation.

Results

As a review, the paper does not present new experimental results. However, it cites key benchmarks and case studies from the literature:

  • Virtual screening of ultra-large libraries (e.g., >1 billion compounds) can achieve hit rates of 10-30% in prospective studies, compared to <1% for traditional high-throughput screening.
  • Deep learning models for binding affinity prediction have reached Pearson correlations of 0.7-0.8 on standard benchmarks (e.g., PDBbind, DUD-E), though performance varies by target.
  • Iterative screening approaches can reduce the number of compounds to be docked by 10-100x while maintaining similar hit rates.

Significance

The broader impact on AI is substantial. This paper signals that drug discovery is becoming a major application domain for machine learning, driving demand for new architectures (e.g., graph neural networks, transformers for molecular generation), large-scale training datasets, and efficient inference methods. It also highlights the need for robust evaluation protocols to avoid overfitting and ensure generalization to novel targets. For AI researchers, the challenges outlined—such as handling chemical space sparsity, incorporating 3D geometry, and achieving interpretability—represent fertile ground for innovation. Ultimately, the review positions computational drug discovery as a high-impact, data-rich field where AI can directly contribute to human health.