A Reproducible and Explainable Deep Learning Framework for Dermoscopic Image Analysis in Skin Cancer
Background/Objectives: Cutaneous melanoma caused over 331,000 new cases and 58,000 deaths worldwide in 2022, with five-year survival falling from 99.4% in localised disease to 35.6% after distant metastasis. Most dermoscopic deep learning studies report headline accuracy without addressing data leakage, calibration, or how their comparisons are constructed. This study asks how much of that performance survives a protocol designed to remove the optimism. Materials and Methods: Six classification architectures were trained on HAM10000/ISIC 2018 (10,015 images, seven classes) at a fixed resolution and with fixed ImageNet-1k pretraining, the architecture confounding with neither. Five segmentation networks were trained alongside MedSAM and evaluated separately under reference-derived and fully automatic box prompts. Lesion-level StratifiedGroupKFold prevented duplicate-lesion leakage; temperature scaling used an inner-validation partition, never the test fold. Models were compared across outer folds (N = 5), saliency localisation against a fixed centred-disc baseline, and generalisation on BCN20000 (15,172 images). Results: The six configurations were statistically indistinguishable (Friedman χ2 = 7.057, p = 0.216; balanced accuracy 0.664–0.700), so no best architecture could be identified. MedSAM ranked first with reference-derived prompts (Dice 0.9547) and last with automatic prompts (0.7729), 16.7 points below a plain U-Net, its advantage reflecting prompt provenance rather than the model. All three class activation methods scored below the fixed disc on mask-IoU (0.2531–0.3512 against 0.4745). Externally, balanced accuracy fell 25 points and the ranking reversed, while AUC held at 0.80–0.84; recalibration helped only two of six models. Conclusions: Under a protocol that isolates each compared factor, the differences that such studies report do not survive. Cross-cohort loss is large and not reliably repaired by recalibration; the leverage lies in evaluation methodology, not architectural novelty.
Authors
- Amal Khalifa Alkhalifa (ORCID: https://orcid.org/0000-0002-7273-4041)
- Sarah A. Alzakari (ORCID: https://orcid.org/0000-0001-8265-2421)
- Abedelmalek Kalefh Tabnjh (ORCID: https://orcid.org/0000-0001-8263-4944)
- Cemil Çolak (ORCID: https://orcid.org/0000-0001-5406-098X)
- Fatma Hilal Yağın (ORCID: https://orcid.org/0000-0002-9848-7958)
- Burak Yagin (ORCID: https://orcid.org/0000-0001-6687-979X)
- Abdulvahap Pınar (ORCID: https://orcid.org/0000-0002-3662-2579)
Institutions
- Princess Nourah bint Abdulrahman University (SA)
- University of Turku (FI)
- Jordan University of Science and Technology (JO)
- Inonu University (TR)
- Turgut Özal University (TR)
- Saveetha University (IN)
- University of Gothenburg (SE)
Publication Details
- Journal
- Journal of Clinical Medicine
- Published
- 2026-09-24
- DOI
- https://doi.org/10.3390/jcm15197422
- Primary Topic
- Cutaneous Melanoma Detection and Management
- Type
- article
- Field-Weighted Citation Impact
- 0.00