Beyond Internal Performance: Evaluating Loss, Augmentation, Fusion, and Generalization in Skin Lesion Classification
Skin lesion classification studies often evaluate a single training configuration using aggregate performance measures that can obscure minority-class failure. This study presents a systematic evaluation of loss functions, preprocessing and augmentation components, test-time augmentation, metadata-fusion strategies, and external generalization on HAM10000 under its natural class distribution. Lesion-grouped splitting is used to prevent images of the same lesion from appearing across training and evaluation sets. The experiments show that the effects of preprocessing and augmentation depend on the underlying loss function, and that model performance varies across fold-trained checkpoints. External evaluation on a filtered seven-class subset of ISIC 2019 reveals substantial degradation under distribution shift. A separate single-run comparison of metadata-fusion strategies shows that the strongest strategy internally does not remain dominant across evaluation criteria on the external dataset. Overall, the findings demonstrate that internal classification performance, calibration, and external generalization capture different aspects of model behaviour. The study highlights the importance of controlled ablations, lesion-aware splitting, multiple evaluation criteria, and external validation when developing classifiers for naturally imbalanced medical-image datasets.
Authors
- Krisha Garg (ORCID: https://orcid.org/0009-0004-3452-0220)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-01
- DOI
- https://doi.org/10.5281/zenodo.22236180
- Primary Topic
- Cutaneous Melanoma Detection and Management
- Type
- preprint