Preprint
Computer Vision

Transparency of deep neural networks for medical image analysis: A review of interpretability methods

Zohaib Salahuddin(Maastricht University), Henry C. Woodruff(Maastricht University Medical Centre), Avishek Chatterjee(Maastricht University), Philippe Lambin(Maastricht University Medical Centre)
December 4, 2021Computers in Biology and Medicine491 citations

491

Citations

8

Influential Citations

Computers in Biology and Medicine

Venue

2021

Year

Abstract

Artificial Intelligence (AI) has emerged as a useful aid in numerous clinical applications for diagnosis and treatment decisions. Deep neural networks have shown the same or better performance than clinicians in many tasks owing to the rapid increase in the available data and computational power. In order to conform to the principles of trustworthy AI, it is essential that the AI system be transparent, robust, fair, and ensure accountability. Current deep neural solutions are referred to as black-boxes due to a lack of understanding of the specifics concerning the decision-making process. Therefore, there is a need to ensure the interpretability of deep neural networks before they can be incorporated into the routine clinical workflow. In this narrative review, we utilized systematic keyword searches and domain expertise to identify nine different types of interpretability methods that have been used for understanding deep learning models for medical image analysis applications based on the type of generated explanations and technical similarities. Furthermore, we report the progress made towards evaluating the explanations produced by various interpretability methods. Finally, we discuss limitations, provide guidelines for using interpretability methods and future directions concerning the interpretability of deep neural networks for medical imaging analysis.

Analysis

Why This Paper Matters

As deep neural networks (DNNs) achieve performance comparable to or exceeding clinicians in medical image analysis, their 'black-box' nature hinders clinical adoption. Trustworthy AI principles—transparency, robustness, fairness, and accountability—demand interpretability. This review addresses a critical gap by systematically organizing the fragmented landscape of interpretability methods, offering a taxonomy that helps researchers and practitioners navigate options. It also highlights the underdeveloped area of evaluating explanations, which is essential for validating that interpretations are faithful and clinically useful.

The paper's significance lies in its practical orientation: it not only categorizes methods but also provides guidelines for their use and discusses limitations. This is particularly valuable for AI practitioners in healthcare, where regulatory and ethical considerations are paramount. By consolidating current knowledge, the review accelerates the path toward integrating interpretable AI into routine clinical workflows.

Technical Contributions

  • Taxonomy of Interpretability Methods: The authors identify nine distinct types of interpretability methods, likely including saliency maps, class activation maps, perturbation-based, gradient-based, and concept-based approaches, among others. This classification is based on the nature of explanations (e.g., local vs. global, feature attribution vs. concept) and technical similarities.
  • Evaluation Framework: The review reports on progress in evaluating explanations, discussing metrics such as faithfulness, sensitivity, and localization accuracy, and emphasizes the need for standardized evaluation protocols.
  • Guidelines for Use: Practical recommendations are provided for selecting appropriate interpretability methods based on the clinical task, model architecture, and desired explanation type.
  • Future Directions: The paper outlines research needs, including developing more robust evaluation metrics, addressing uncertainty, and creating methods that provide causal or counterfactual explanations.

Results

The paper is a narrative review, so it does not present new experimental results. Instead, its 'results' are the categorization of nine interpretability method types and the synthesis of evaluation approaches. The authors report that while many methods exist, there is no consensus on how to evaluate their quality, and many evaluations are qualitative or task-specific. They note that current methods often produce inconsistent or misleading explanations, underscoring the need for standardized validation.

Significance

This review serves as a foundational reference for researchers and clinicians seeking to implement interpretable AI in medical imaging. By providing a clear taxonomy and guidelines, it reduces the barrier to entry and promotes best practices. It also highlights the importance of evaluation, pushing the field toward more rigorous and trustworthy AI systems. The paper's impact extends beyond medical imaging to any high-stakes AI application where transparency is critical, reinforcing the broader movement toward explainable and responsible AI.