Assessing the impact of layer selection on CAM-based explainability for YOLOv8: a study on hand-sketched digital logic circuits
Abstract Class activation mapping (CAM) techniques improve the interpretability of deep neural networks by highlighting image regions that contribute to model predictions. However, the quality of CAM explanations can vary considerably according to the selected target layer. This study presents a layer-wise ablation analysis of seven CAM techniques—Grad-CAM, Grad-CAM++, HiResCAM, XGrad-CAM, LayerCAM, EigenCAM, and EigenGradCAM—applied to a YOLOv8 detector of hand-sketched digital logic circuits. Explanation quality was evaluated quantitatively using Focus Score, Pointing Game hit rate (PG-Hit), intersection over union (IoU), and Sharpness, supported by qualitative inspection of the generated saliency maps. The results demonstrate that explanation quality depends jointly on the CAM technique and target-layer depth. Within the evaluated dataset and experimental setting, LayerCAM and EigenGradCAM applied to the fused mid-deep configuration $$[9,12,15]$$ produced the most semantically meaningful explanations. Although deeper configurations generally incurred greater computational cost, the $$[9,12,15]$$ configuration achieved a mean latency of approximately 15.5 ms. Under the reported hardware and measurement conditions, this value was below the 33.3-ms processing-time requirement corresponding to 30 frames per second.
Authors
- Mohamed Waleed Fakhr (ORCID: https://orcid.org/0000-0001-5147-2639)
- Noha ElMasry
- Fahima A. Maghraby (ORCID: https://orcid.org/0000-0002-2547-5673)
Institutions
- Misr International University (EG)
- Arab Academy for Science, Technology, and Maritime Transport (EG)
Publication Details
- Journal
- Scientific Reports
- Published
- 2026-10-07
- DOI
- https://doi.org/10.1038/s41598-026-71669-x
- Primary Topic
- Explainable Artificial Intelligence (XAI)
- Type
- article
- Field-Weighted Citation Impact
- 0.00