A Reproducible and Explainable Deep Learning Framework for Dermoscopic Image Analysis in Skin Cancer

Background/Objectives: Cutaneous melanoma caused over 331,000 new cases and 58,000 deaths worldwide in 2022, with five-year survival falling from 99.4% in localised disease to 35.6% after distant metastasis. Most dermoscopic deep learning studies report headline accuracy without addressing data leakage, calibration, or how their comparisons are constructed. This study asks how much of that performance survives a protocol designed to remove the optimism. Materials and Methods: Six classification architectures were trained on HAM10000/ISIC 2018 (10,015 images, seven classes) at a fixed resolution and with fixed ImageNet-1k pretraining, the architecture confounding with neither. Five segmentation networks were trained alongside MedSAM and evaluated separately under reference-derived and fully automatic box prompts. Lesion-level StratifiedGroupKFold prevented duplicate-lesion leakage; temperature scaling used an inner-validation partition, never the test fold. Models were compared across outer folds (N = 5), saliency localisation against a fixed centred-disc baseline, and generalisation on BCN20000 (15,172 images). Results: The six configurations were statistically indistinguishable (Friedman χ2 = 7.057, p = 0.216; balanced accuracy 0.664–0.700), so no best architecture could be identified. MedSAM ranked first with reference-derived prompts (Dice 0.9547) and last with automatic prompts (0.7729), 16.7 points below a plain U-Net, its advantage reflecting prompt provenance rather than the model. All three class activation methods scored below the fixed disc on mask-IoU (0.2531–0.3512 against 0.4745). Externally, balanced accuracy fell 25 points and the ranking reversed, while AUC held at 0.80–0.84; recalibration helped only two of six models. Conclusions: Under a protocol that isolates each compared factor, the differences that such studies report do not survive. Cross-cohort loss is large and not reliably repaired by recalibration; the leverage lies in evaluation methodology, not architectural novelty.

Authors

Institutions

Publication Details

Journal
Journal of Clinical Medicine
Published
2026-09-24
DOI
https://doi.org/10.3390/jcm15197422
Primary Topic
Cutaneous Melanoma Detection and Management
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

A Reproducible and Explainable Deep Learning Framework for Dermoscopic Image Analysis in Skin Cancer

Amal Khalifa Alkhalifa, Sarah A. Alzakari, Abedelmalek Kalefh Tabnjh, Cemil Çolak et al.
Journal of Clinical Medicine
Cutaneous Melanoma Detection and Management
article

A Reproducible and Explainable Deep Learning Framework for Dermoscopic Image Analysis in Skin Cancer

Amal Khalifa Alkhalifa, Sarah A. Alzakari, Abedelmalek Kalefh Tabnjh, Cemil Çolak, Fatma Hilal Yağın, Burak Yagin, Abdulvahap Pınar
article en

Abstract

Background/Objectives: Cutaneous melanoma caused over 331,000 new cases and 58,000 deaths worldwide in 2022, with five-year survival falling from 99.4% in localised disease to 35.6% after distant metastasis. Most dermoscopic deep learning studies report headline accuracy without addressing data leakage, calibration, or how their comparisons are constructed. This study asks how much of that performance survives a protocol designed to remove the optimism. Materials and Methods: Six classification architectures were trained on HAM10000/ISIC 2018 (10,015 images, seven classes) at a fixed resolution and with fixed ImageNet-1k pretraining, the architecture confounding with neither. Five segmentation networks were trained alongside MedSAM and evaluated separately under reference-derived and fully automatic box prompts. Lesion-level StratifiedGroupKFold prevented duplicate-lesion leakage; temperature scaling used an inner-validation partition, never the test fold. Models were compared across outer folds (N = 5), saliency localisation against a fixed centred-disc baseline, and generalisation on BCN20000 (15,172 images). Results: The six configurations were statistically indistinguishable (Friedman χ2 = 7.057, p = 0.216; balanced accuracy 0.664–0.700), so no best architecture could be identified. MedSAM ranked first with reference-derived prompts (Dice 0.9547) and last with automatic prompts (0.7729), 16.7 points below a plain U-Net, its advantage reflecting prompt provenance rather than the model. All three class activation methods scored below the fixed disc on mask-IoU (0.2531–0.3512 against 0.4745). Externally, balanced accuracy fell 25 points and the ranking reversed, while AUC held at 0.80–0.84; recalibration helped only two of six models. Conclusions: Under a protocol that isolates each compared factor, the differences that such studies report do not survive. Cross-cohort loss is large and not reliably repaired by recalibration; the leverage lies in evaluation methodology, not architectural novelty.

Journal of Clinical MedicineVol. 15(19)
Princess Nourah bint Abdulrahman University (SA), University of Turku (FI), Jordan University of Science and Technology (JO), Inonu University (TR), Turgut Özal University (TR), Saveetha University (IN), University of Gothenburg (SE)
Openalex Percentile: Top 14%
Cutaneous Melanoma Detection and Management
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.