Interpretable hybrid model combining CNN-ViT with Grad-CAM and SHAP for transparent mapping of land use changes

Abstract Accurate mapping of land use and land cover (LULC) in arid and rapidly urbanizing environments remains challenging due to high spectral similarity among land-cover classes and the inherent topographic and environmental complexity of these regions. This study introduces a novel deep learning (DL)–based framework for multi-temporal LULC classification of Al-Taif City, Saudi Arabia, using Landsat 8 imagery from 2013 to 2024. The proposed architecture integrates three complementary DL models: InceptionV3 for multi-scale feature extraction, EfficientNet for efficient and discriminative representation learning, and Vision Transformer (ViT) for capturing long-range contextual dependencies. The feature representations generated by these models are concatenated and subsequently fused to construct a comprehensive feature space. To reduce redundancy and improve feature discriminability, Principal Component Analysis (PCA) is applied to the fused representations prior to classification, which incorporates both spectral indices and learned deep features. To enhance interpretability and transparency, the proposed framework employs Gradient-weighted Class Activation Mapping (Grad-CAM) to provide spatial visual explanations of model predictions, along with SHapley Additive exPlanations (SHAP) to quantify the contribution of individual input features to the final classification outcomes. Experimental results demonstrate the robustness and high performance of the proposed approach, achieving overall accuracies of 97.79% (κ = 0.967) for 2013 and 98.87% (κ = 0.985) for 2024. Furthermore, the temporal analysis reveals significant LULC dynamics within the study area, characterized by a marked expansion of urban and industrial zones alongside a noticeable decline in agricultural land. These findings highlight the effectiveness of the proposed framework in delivering both high-precision classification and interpretable insights, making it a reliable tool for long-term LULC monitoring and sustainable urban planning in arid regions.

Authors

Institutions

Publication Details

Journal
Scientific Reports
Published
2026-09-21
DOI
https://doi.org/10.1038/s41598-026-72111-y
Primary Topic
Remote-Sensing Image Classification
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Interpretable hybrid model combining CNN-ViT with Grad-CAM and SHAP for transparent mapping of land use changes

Tawfiq Hasanin, Abdullah Saad AL-Malaise AL-Ghamdi, Abdulmajid A. Alnoamani
Scientific Reports
Remote-Sensing Image Classification
article

Interpretable hybrid model combining CNN-ViT with Grad-CAM and SHAP for transparent mapping of land use changes

Tawfiq Hasanin, Abdullah Saad AL-Malaise AL-Ghamdi, Abdulmajid A. Alnoamani
article en

Abstract

Abstract Accurate mapping of land use and land cover (LULC) in arid and rapidly urbanizing environments remains challenging due to high spectral similarity among land-cover classes and the inherent topographic and environmental complexity of these regions. This study introduces a novel deep learning (DL)–based framework for multi-temporal LULC classification of Al-Taif City, Saudi Arabia, using Landsat 8 imagery from 2013 to 2024. The proposed architecture integrates three complementary DL models: InceptionV3 for multi-scale feature extraction, EfficientNet for efficient and discriminative representation learning, and Vision Transformer (ViT) for capturing long-range contextual dependencies. The feature representations generated by these models are concatenated and subsequently fused to construct a comprehensive feature space. To reduce redundancy and improve feature discriminability, Principal Component Analysis (PCA) is applied to the fused representations prior to classification, which incorporates both spectral indices and learned deep features. To enhance interpretability and transparency, the proposed framework employs Gradient-weighted Class Activation Mapping (Grad-CAM) to provide spatial visual explanations of model predictions, along with SHapley Additive exPlanations (SHAP) to quantify the contribution of individual input features to the final classification outcomes. Experimental results demonstrate the robustness and high performance of the proposed approach, achieving overall accuracies of 97.79% (κ = 0.967) for 2013 and 98.87% (κ = 0.985) for 2024. Furthermore, the temporal analysis reveals significant LULC dynamics within the study area, characterized by a marked expansion of urban and industrial zones alongside a noticeable decline in agricultural land. These findings highlight the effectiveness of the proposed framework in delivering both high-precision classification and interpretable insights, making it a reliable tool for long-term LULC monitoring and sustainable urban planning in arid regions.

Scientific Reports
King Abdulaziz University (SA)
Sustainable cities and communities
Openalex Percentile: Top 14%
Remote-Sensing Image Classification
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Interpretable hybrid model combining CNN-ViT with Grad-CAM and SHAP for transparent mapping of land use changes — Tawfiq Hasanin, Abdullah Saad AL-Malaise AL-Ghamdi, et al. · Scientific Reports (2026) | TGRS Research Map | TGRS