Preprint
AI Safety & Alignment

Interpreting Black-Box Models: A Review on Explainable Artificial Intelligence

Vikas Hassija(KIIT University), Vinay Chamola(Birla Institute of Technology and Science, Pilani), Atmesh Mahapatra(Birla Institute of Technology and Science, Pilani), Abhinandan Singal(Jaypee Institute of Information Technology), Divyansh Goel(Jaypee Institute of Information Technology), Kaizhu Huang(Duke Kunshan University), Simone Scardapane(Sapienza University of Rome), Indro Spinelli(Istituto Nazionale di Fisica Nucleare, Sezione di Roma I), Mufti Mahmud(Nottingham Trent University), Amir Hussain(Edinburgh Napier University)
August 24, 2023Cognitive Computation1,825 citations

1.8k

Citations

27

Influential Citations

Cognitive Computation

Venue

2023

Year

Abstract

Abstract Recent years have seen a tremendous growth in Artificial Intelligence (AI)-based methodological development in a broad range of domains. In this rapidly evolving field, large number of methods are being reported using machine learning (ML) and Deep Learning (DL) models. Majority of these models are inherently complex and lacks explanations of the decision making process causing these models to be termed as 'Black-Box'. One of the major bottlenecks to adopt such models in mission-critical application domains, such as banking, e-commerce, healthcare, and public services and safety, is the difficulty in interpreting them. Due to the rapid proleferation of these AI models, explaining their learning and decision making process are getting harder which require transparency and easy predictability. Aiming to collate the current state-of-the-art in interpreting the black-box models, this study provides a comprehensive analysis of the explainable AI (XAI) models. To reduce false negative and false positive outcomes of these back-box models, finding flaws in them is still difficult and inefficient. In this paper, the development of XAI is reviewed meticulously through careful selection and analysis of the current state-of-the-art of XAI research. It also provides a comprehensive and in-depth evaluation of the XAI frameworks and their efficacy to serve as a starting point of XAI for applied and theoretical researchers. Towards the end, it highlights emerging and critical issues pertaining to XAI research to showcase major, model-specific trends for better explanation, enhanced transparency, and improved prediction accuracy.

Analysis

Why This Paper Matters

This paper addresses a critical bottleneck in the adoption of AI systems in high-stakes domains like healthcare, banking, and public safety: the lack of interpretability of black-box models. As AI models become more complex, their decision-making processes grow opaque, hindering trust and regulatory compliance. By providing a comprehensive review of explainable AI (XAI) methods, this work equips practitioners with a structured understanding of available techniques, from model-specific to model-agnostic approaches. The paper's high citation count (1825) underscores its relevance and utility as a foundational reference in the rapidly growing XAI field.

The review is particularly timely given the increasing regulatory pressure for algorithmic transparency (e.g., GDPR's right to explanation). It systematically categorizes XAI frameworks, evaluates their efficacy, and highlights unresolved challenges, such as the difficulty in detecting and reducing false negatives/positives. This makes it a valuable resource for researchers and engineers seeking to deploy trustworthy AI systems.

Technical Contributions

  • Comprehensive Literature Review: The paper meticulously surveys the state-of-the-art in XAI, covering a broad range of methods including LIME, SHAP, Grad-CAM, and attention-based explanations.
  • Framework Evaluation: It provides an in-depth evaluation of XAI frameworks, assessing their strengths and weaknesses for different use cases (e.g., image classification, natural language processing).
  • Identification of Trends: The review highlights model-specific trends for improving explanation quality, transparency, and prediction accuracy, offering guidance for future research.
  • Critical Issue Analysis: It discusses emerging issues such as the trade-off between accuracy and interpretability, and the challenge of evaluating explanation fidelity.

Results

The paper does not present new experimental results but synthesizes findings from the literature. It notes that current XAI methods still struggle with efficiently finding flaws in black-box models to reduce false negatives and false positives. The review concludes that while progress has been made, significant gaps remain in achieving both high accuracy and full transparency.

Significance

This review has broad impact by serving as a one-stop reference for XAI research, lowering the barrier for entry for applied researchers and practitioners. It helps bridge the gap between complex AI models and their deployment in mission-critical applications, potentially accelerating the responsible adoption of AI. The identified trends and open issues also guide future research directions, making it a seminal work in the AI safety and alignment landscape.