Attention-aligned guided gradient-based class activation mapping for explainable indian sign language recognition
Sign language recognition is an important assistive technology for improving communication accessibility for individuals with hearing and speech impairments. While recent deep learning models have achieved impressive recognition accuracy, they often operate as ”black boxes”. This lack of transparency limits their trustworthiness and adoption in real-world assistive technologies. To address this, we introduce an explainability-aligned hybrid framework that combines Convolutional Neural Networks (CNNs) and Transformers for interpretable Indian Sign Language (ISL) recognition. Our architecture pairs an EfficientNet-based convolutional branch, designed to capture fine-grained, local hand features with a Shifted Window Transformer (Swin Transformer) branch that models global contextual relationships. Further, to ensure the model’s decisions make intuitive sense, we introduce a learnable explanation-alignment mechanism. This mechanism uses an explanation-alignment loss and an adaptive gating strategy to enforce spatial consistency between the CNN’s saliency maps and the Transformer’s attention representations. To evaluate the proposed approach, a new Indian Sign Language dataset containing 82,500 images from 33 sign classes was developed and assessed using a subject-independent evaluation protocol. The proposed framework achieved 99.85 ± 0.06 % classification accuracy and 99.99 ± 0.05 % Macro-F1 (macro-averaged F1-score), outperforming standalone Convolutional Neural Network, Transformer, and conventional hybrid architectures. Furthermore, explainability assessments using deletion faithfulness and weakly supervised object localization metrics confirm that our approach generates more reliable, well-focused, and spatially coherent visual explanations. The findings suggest that incorporating explanation alignment into hybrid learning architectures can simultaneously improve recognition accuracy and explanation quality, providing a promising step toward more trustworthy and interpretable Artificial Intelligence systems for sign language recognition.
Authors
- Raman Maini
- Sandhya Rani
- Kavita Gupta
- Rajeev Goel
Institutions
- Punjabi University (IN)
Publication Details
- Journal
- Engineering Applications of Artificial Intelligence
- Published
- 2026-09-30
- DOI
- https://doi.org/10.1016/j.engappai.2026.116149
- Primary Topic
- Hand Gesture Recognition Systems
- Type
- article
- Field-Weighted Citation Impact
- 0.00