Explainable Artificial Intelligence for Computer Vision: A Comprehensive Study on Interpretability Techniques, and Applications
Abstract Convolutional neural networks (CNNs) have achieved state-of-the-art performance across computer vision tasks, including image classification, object detection, and medical image analysis. However, the opaque, black-box nature of these models limits their adoption in high-stakes domains such as healthcare, autonomous driving, and security, where accountability and trust are essential. Explainable Artificial Intelligence (XAI) has emerged as a critical research direction that seeks to make the decision-making process of deep models transparent and interpretable to human users without materially sacrificing predictive performance. While prior work has largely evaluated individual explainability techniques in isolation, few studies jointly deploy and quantitatively compare complementary explanation modalities within a single inference pipeline. We review recent literature published between 2023 and 2026, categorize explainability methods along the model-agnostic/model-specific and global/local dimensions, and provide a detailed technical analysis of five widely used techniques: Local Interpretable Model-Agnostic Explanations (LIME), SHapley Additive exPlanations (SHAP), Gradient-weighted Class Activation Mapping (Grad-CAM), Integrated Gradients, and saliency/attention-based maps. Building on this analysis, we propose an XAI-integrated computer vision framework that couples a fine-tuned Residual Network (ResNet-18) classifier with Grad-CAM, SHAP, and LIME explanation modules, and introduce an Explainability Consensus Score that quantifies the spatial agreement among the three explanation modalities for a given prediction. The framework is designed and specified for evaluation on the Chest X-ray Pneumonia dataset, with a complete evaluation protocol, planned performance metrics, and explainability-agreement metrics defined in advance of full experimental deployment; quantitative results will be reported following implementation and are intentionally not fabricated in this manuscript. Based on the reviewed literature and the qualitative behavior of the proposed explanation modules, we argue that combining local (SHAP, LIME) and gradient-based (Grad-CAM) explanations is expected to improve diagnostic transparency over any single technique used in isolation, a hypothesis the proposed Consensus Score is designed to test quantitatively. The paper concludes by discussing open challenges, including explanation faithfulness, computational overhead, bias propagation, and the evaluation of explainable vision transformers, and outlines future research directions toward standardized, responsible, and human-centered explainable computer vision systems. Keywords: Explainable Artificial Intelligence, Computer Vision, Deep Learning, Convolutional Neural Network, Grad-CAM, SHAP, LIME, Explainability Consensus Score, Model Interpretability, Trustworthy AI.
Authors
- Vishal Shrivastava (ORCID: https://orcid.org/0000-0002-8353-2752)
- Akhil Pandey
- Gaurav Jangid
- Ishan Gupta
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-05
- DOI
- https://doi.org/10.5281/zenodo.23160759
- Primary Topic
- Explainable Artificial Intelligence (XAI)
- Type
- article
- Field-Weighted Citation Impact
- 0.00