Attention-aligned guided gradient-based class activation mapping for explainable indian sign language recognition

Sign language recognition is an important assistive technology for improving communication accessibility for individuals with hearing and speech impairments. While recent deep learning models have achieved impressive recognition accuracy, they often operate as ”black boxes”. This lack of transparency limits their trustworthiness and adoption in real-world assistive technologies. To address this, we introduce an explainability-aligned hybrid framework that combines Convolutional Neural Networks (CNNs) and Transformers for interpretable Indian Sign Language (ISL) recognition. Our architecture pairs an EfficientNet-based convolutional branch, designed to capture fine-grained, local hand features with a Shifted Window Transformer (Swin Transformer) branch that models global contextual relationships. Further, to ensure the model’s decisions make intuitive sense, we introduce a learnable explanation-alignment mechanism. This mechanism uses an explanation-alignment loss and an adaptive gating strategy to enforce spatial consistency between the CNN’s saliency maps and the Transformer’s attention representations. To evaluate the proposed approach, a new Indian Sign Language dataset containing 82,500 images from 33 sign classes was developed and assessed using a subject-independent evaluation protocol. The proposed framework achieved 99.85 ± 0.06 % classification accuracy and 99.99 ± 0.05 % Macro-F1 (macro-averaged F1-score), outperforming standalone Convolutional Neural Network, Transformer, and conventional hybrid architectures. Furthermore, explainability assessments using deletion faithfulness and weakly supervised object localization metrics confirm that our approach generates more reliable, well-focused, and spatially coherent visual explanations. The findings suggest that incorporating explanation alignment into hybrid learning architectures can simultaneously improve recognition accuracy and explanation quality, providing a promising step toward more trustworthy and interpretable Artificial Intelligence systems for sign language recognition.

Authors

Institutions

Publication Details

Journal
Engineering Applications of Artificial Intelligence
Published
2026-09-30
DOI
https://doi.org/10.1016/j.engappai.2026.116149
Primary Topic
Hand Gesture Recognition Systems
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Attention-aligned guided gradient-based class activation mapping for explainable indian sign language recognition

Raman Maini, Sandhya Rani, Kavita Gupta, Rajeev Goel
Engineering Applications of Artificial Intelligence
Hand Gesture Recognition Systems
article

Attention-aligned guided gradient-based class activation mapping for explainable indian sign language recognition

Raman Maini, Sandhya Rani, Kavita Gupta, Rajeev Goel
article en

Abstract

Sign language recognition is an important assistive technology for improving communication accessibility for individuals with hearing and speech impairments. While recent deep learning models have achieved impressive recognition accuracy, they often operate as ”black boxes”. This lack of transparency limits their trustworthiness and adoption in real-world assistive technologies. To address this, we introduce an explainability-aligned hybrid framework that combines Convolutional Neural Networks (CNNs) and Transformers for interpretable Indian Sign Language (ISL) recognition. Our architecture pairs an EfficientNet-based convolutional branch, designed to capture fine-grained, local hand features with a Shifted Window Transformer (Swin Transformer) branch that models global contextual relationships. Further, to ensure the model’s decisions make intuitive sense, we introduce a learnable explanation-alignment mechanism. This mechanism uses an explanation-alignment loss and an adaptive gating strategy to enforce spatial consistency between the CNN’s saliency maps and the Transformer’s attention representations. To evaluate the proposed approach, a new Indian Sign Language dataset containing 82,500 images from 33 sign classes was developed and assessed using a subject-independent evaluation protocol. The proposed framework achieved 99.85 ± 0.06 % classification accuracy and 99.99 ± 0.05 % Macro-F1 (macro-averaged F1-score), outperforming standalone Convolutional Neural Network, Transformer, and conventional hybrid architectures. Furthermore, explainability assessments using deletion faithfulness and weakly supervised object localization metrics confirm that our approach generates more reliable, well-focused, and spatially coherent visual explanations. The findings suggest that incorporating explanation alignment into hybrid learning architectures can simultaneously improve recognition accuracy and explanation quality, providing a promising step toward more trustworthy and interpretable Artificial Intelligence systems for sign language recognition.

Engineering Applications of Artificial IntelligenceVol. 184
Punjabi University (IN)
Quality Education
Openalex Percentile: Top 9%
Hand Gesture Recognition Systems
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.