Journal Article
Machine Learning

Explainable Machine Learning for Scientific Insights and Discoveries

Ribana Roscher(University of Bonn), Bastian Bohn(University of Bonn), Marco F. Duarte(University of Massachusetts Amherst), Jochen Garcke(University of Bonn)
January 1, 2020IEEE Access986 citations

986

Citations

29

Influential Citations

IEEE Access

Venue

2020

Year

Abstract

Machine learning methods have been remarkably successful for a wide range of application areas in the extraction of essential information from data. An exciting and relatively recent development is the uptake of machine learning in the natural sciences, where the major goal is to obtain novel scientific insights and discoveries from observational or simulated data. A prerequisite for obtaining a scientific outcome is domain knowledge, which is needed to gain explainability, but also to enhance scientific consistency. In this article, we review explainable machine learning in view of applications in the natural sciences and discuss three core elements that we identified as relevant in this context: transparency, interpretability, and explainability. With respect to these core elements, we provide a survey of recent scientific works that incorporate machine learning and the way that explainable machine learning is used in combination with domain knowledge from the application areas.

Analysis

Why This Paper Matters

This paper addresses a critical gap in the application of machine learning to natural sciences: the need for explainability to derive scientific insights rather than just predictions. As ML models become more complex, their black-box nature hinders adoption in fields where understanding mechanisms is paramount. By framing explainability through transparency, interpretability, and explainability, the authors provide a taxonomy that helps practitioners choose appropriate methods for their domain.

The survey is timely given the surge of ML in scientific research, from drug discovery to climate modeling. It emphasizes that domain knowledge is not just a supplement but a prerequisite for meaningful scientific outcomes, challenging the notion that purely data-driven approaches suffice. This perspective is valuable for AI practitioners seeking to bridge the gap between predictive performance and scientific rigor.

Technical Contributions

  • Core Elements Framework: Defines transparency (how model works), interpretability (how decisions are made), and explainability (why decisions are made) as distinct but interrelated concepts.
  • Domain Knowledge Integration: Shows how prior scientific knowledge can be embedded into ML pipelines to enhance consistency and explainability.
  • Survey of Applications: Covers examples from astronomy, biology, and materials science where explainable ML led to new discoveries.
  • Method Categorization: Groups explainability methods into intrinsic (e.g., linear models) and post-hoc (e.g., LIME, SHAP) approaches, with guidance on suitability for scientific tasks.

Results

As a review paper, no quantitative results are presented. However, the paper synthesizes findings from over 100 cited works, demonstrating that explainable ML methods have successfully identified new physical laws, discovered biomarkers, and optimized experimental designs. The framework itself is a qualitative contribution that organizes a rapidly growing field.

Significance

This paper has influenced subsequent research by providing a common vocabulary for explainability in science. It has been cited nearly 1000 times, indicating its impact on both ML and scientific communities. For AI practitioners, it underscores that building interpretable models is not a trade-off but a necessity for scientific credibility. The emphasis on domain knowledge also encourages interdisciplinary collaboration, which is essential for real-world impact.