From Saliency Maps to Circuits: An Integrative Survey of Explainable AI and Mechanistic Interpretability

Artificial intelligence (AI) models, particularly those based on deep learning, have achieved strong performance across a wide range of real-world applications. However, as models increase in scale and complexity, their behavior and internal decision-making processes become increasingly difficult to understand. Explainable artificial intelligence (XAI) has emerged to address this challenge by developing methods for making model predictions and internal mechanisms more transparent and intelligible. Despite substantial progress, the XAI landscape remains difficult to navigate due to methodological diversity, inconsistent terminology, and increasingly distinct research directions. In particular, classical XAI techniques, which primarily explain model predictions and behavior, and mechanistic interpretability (MI), which seeks to understand internal representations and computations, are often studied separately. This paper provides an integrative survey of these approaches, bringing them together within a shared taxonomy that spans explanatory scope, stage, and methodological foundations. The review examines their underlying principles, assumptions, strengths, and limitations while also surveying evaluation criteria for assessing explanation quality and highlighting connections and distinctions across the broader interpretability landscape. By synthesizing these research directions within a common framework, the study aims to provide a cohesive reference for understanding, organizing, and comparing methods for explaining machine learning models.

Authors

Institutions

Publication Details

Journal
Machine Learning and Knowledge Extraction
Published
2026-10-05
DOI
https://doi.org/10.3390/make8100315
Primary Topic
Explainable Artificial Intelligence (XAI)
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

From Saliency Maps to Circuits: An Integrative Survey of Explainable AI and Mechanistic Interpretability

Adina Magda Florea, Andrei Dugăeșescu
Machine Learning and Knowledge Extraction
Explainable Artificial Intelligence (XAI)
article

From Saliency Maps to Circuits: An Integrative Survey of Explainable AI and Mechanistic Interpretability

Adina Magda Florea, Andrei Dugăeșescu
article en

Abstract

Artificial intelligence (AI) models, particularly those based on deep learning, have achieved strong performance across a wide range of real-world applications. However, as models increase in scale and complexity, their behavior and internal decision-making processes become increasingly difficult to understand. Explainable artificial intelligence (XAI) has emerged to address this challenge by developing methods for making model predictions and internal mechanisms more transparent and intelligible. Despite substantial progress, the XAI landscape remains difficult to navigate due to methodological diversity, inconsistent terminology, and increasingly distinct research directions. In particular, classical XAI techniques, which primarily explain model predictions and behavior, and mechanistic interpretability (MI), which seeks to understand internal representations and computations, are often studied separately. This paper provides an integrative survey of these approaches, bringing them together within a shared taxonomy that spans explanatory scope, stage, and methodological foundations. The review examines their underlying principles, assumptions, strengths, and limitations while also surveying evaluation criteria for assessing explanation quality and highlighting connections and distinctions across the broader interpretability landscape. By synthesizing these research directions within a common framework, the study aims to provide a cohesive reference for understanding, organizing, and comparing methods for explaining machine learning models.

Machine Learning and Knowledge ExtractionVol. 8(10)
Universitatea Națională de Știință și Tehnologie Politehnica București (RO)
Openalex Percentile: Top 10%
Explainable Artificial Intelligence (XAI)
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

From Saliency Maps to Circuits: An Integrative Survey of Explainable AI and Mechanistic Interpretability — Adina Magda Florea, Andrei Dugăeșescu · Machine Learning and Knowledge Extraction (2026) | TGRS Research Map | TGRS