Preprint
Machine Learning

An Introduction to Machine Learning

Solveig Badillo(Roche (Switzerland)), Balázs Bánfai(Roche (Switzerland)), Fabian Birzele(Roche (Switzerland)), Iakov I. Davydov(Roche (Switzerland)), Lucy Hutchinson(Roche (Switzerland)), Tony Kam‐Thong(Roche (Switzerland)), Juliane Siebourg‐Polster(Roche (Switzerland)), Bernhard Steiert(Roche (Switzerland)), Jitao David Zhang(Roche (Switzerland))
March 3, 2020Clinical Pharmacology & Therapeutics795 citations

795

Citations

6

Influential Citations

Clinical Pharmacology & Therapeutics

Venue

2020

Year

Abstract

In the last few years, machine learning (ML) and artificial intelligence have seen a new wave of publicity fueled by the huge and ever-increasing amount of data and computational power as well as the discovery of improved learning algorithms. However, the idea of a computer learning some abstract concept from data and applying them to yet unseen situations is not new and has been around at least since the 1950s. Many of these basic principles are very familiar to the pharmacometrics and clinical pharmacology community. In this paper, we want to introduce the foundational ideas of ML to this community such that readers obtain the essential tools they need to understand publications on the topic. Although we will not go into the very details and theoretical background, we aim to point readers to relevant literature and put applications of ML in molecular biology as well as the fields of pharmacometrics and clinical pharmacology into perspective.

Analysis

Why This Paper Matters

This paper addresses a critical gap in the pharmacometrics and clinical pharmacology community: the need for a foundational understanding of machine learning (ML) principles. As ML increasingly permeates drug development—from biomarker discovery to dose optimization—many domain experts lack the vocabulary and conceptual framework to critically evaluate or apply these methods. By presenting ML concepts in familiar terms (e.g., linking regression to pharmacokinetic modeling), the authors lower the barrier to entry for a non-computer-science audience.

The timing of this review (2020) coincides with a surge in ML-driven drug discovery efforts, making it a timely resource. Its high citation count (795) underscores its utility as a reference for both newcomers and educators. The paper’s emphasis on connecting classical ML (e.g., decision trees, SVMs) to domain-specific problems ensures relevance even as deep learning evolves.

Technical Contributions

  • Conceptual mapping: Translates ML jargon (e.g., overfitting, cross-validation) into pharmacometrics analogies, such as comparing regularization to prior distributions in Bayesian modeling.
  • Application taxonomy: Categorizes ML use cases in molecular biology (e.g., protein structure prediction, gene expression analysis) and clinical pharmacology (e.g., patient stratification, adverse event prediction).
  • Literature roadmap: Curates key references for readers who wish to dive deeper into specific algorithms or theoretical foundations.
  • Perspective on limitations: Discusses challenges like data scarcity, interpretability, and regulatory acceptance in clinical settings.

Results

As a review paper, no quantitative results are presented. The primary outcome is the conceptual framework provided to readers, enabling them to better understand and critique ML-based publications in their field. The paper’s impact is measured by its citation count (795) and its role as a pedagogical tool in workshops and courses.

Significance

This paper contributes to the democratization of ML knowledge within a specialized domain. By equipping pharmacometricians and clinical pharmacologists with ML literacy, it facilitates interdisciplinary collaboration and accelerates the integration of AI into drug development pipelines. Its emphasis on foundational concepts (rather than cutting-edge techniques) ensures long-term relevance, as the core principles of ML remain stable even as specific algorithms evolve. The paper also implicitly advocates for rigorous evaluation and validation of ML models in clinical contexts, a message that aligns with broader efforts to improve reproducibility in AI-driven science.