From Saliency Maps to Circuits: An Integrative Survey of Explainable AI and Mechanistic Interpretability
Artificial intelligence (AI) models, particularly those based on deep learning, have achieved strong performance across a wide range of real-world applications. However, as models increase in scale and complexity, their behavior and internal decision-making processes become increasingly difficult to understand. Explainable artificial intelligence (XAI) has emerged to address this challenge by developing methods for making model predictions and internal mechanisms more transparent and intelligible. Despite substantial progress, the XAI landscape remains difficult to navigate due to methodological diversity, inconsistent terminology, and increasingly distinct research directions. In particular, classical XAI techniques, which primarily explain model predictions and behavior, and mechanistic interpretability (MI), which seeks to understand internal representations and computations, are often studied separately. This paper provides an integrative survey of these approaches, bringing them together within a shared taxonomy that spans explanatory scope, stage, and methodological foundations. The review examines their underlying principles, assumptions, strengths, and limitations while also surveying evaluation criteria for assessing explanation quality and highlighting connections and distinctions across the broader interpretability landscape. By synthesizing these research directions within a common framework, the study aims to provide a cohesive reference for understanding, organizing, and comparing methods for explaining machine learning models.
Authors
- Adina Magda Florea (ORCID: https://orcid.org/0000-0001-7249-1871)
- Andrei Dugăeșescu
Institutions
- Universitatea Națională de Știință și Tehnologie Politehnica București (RO)
Publication Details
- Journal
- Machine Learning and Knowledge Extraction
- Published
- 2026-10-05
- DOI
- https://doi.org/10.3390/make8100315
- Primary Topic
- Explainable Artificial Intelligence (XAI)
- Type
- article
- Field-Weighted Citation Impact
- 0.00