Multimodal Swin Transformer with Hyper-Feature and Attention-Gated Clinical Fusion for Skin Lesion Classification

Clinical photographs and patient metadata are complementary sources of information for automated skin lesion classification. In this study, a multimodal framework is proposed that combines Swin Transformer Tiny (Swin-T), Hyper-Feature Fusion (HFF), and Attention-Gated Clinical Metadata Fusion (AGCMF). HFF combines four levels of hierarchical representations from the Swin Transformer into a 256-dimensional visual embedding. A multilayer perceptron encodes 20 clinical metadata variables into a 64-dimensional representation. The features are concatenated, adaptively gated and projected for six-class classification. Five-fold patient-grouped cross-validation was performed on PAD-UFES-20, which consists of 2298 clinical photographs of 1373 patients. The proposed model obtained 83.55% accuracy without test-time augmentation (TTA), 83.29% with TTA, 82.92% weighted F1, 77.29% macro F1, 75.75% balanced accuracy, 93.12% macro-ROC-AUC, 94.59% weighted ROC-AUC, and Cohen’s kappa of 0.7702. In the ablation experiments, HFF increased the accuracy of the Swin-T model from 78.64% to 80.91%. Clinical feature concatenation increased it to 81.76%, and the complete HFF + AGCMF framework achieved 83.29%. These results indicate that hierarchical visual fusion and adaptive multimodal integration provide complementary benefits. The proposed framework achieved 3.89 percentage points higher accuracy than the image-plus-metadata baseline under the same TTA evaluation setting. A biopsy-exclusion sensitivity analysis also revealed that the performance of the model was only slightly affected by the removal of the diagnostic-verification variable, with the 19-feature model maintaining 82.91% accuracy, 82.53% weighted F1, and 94.21% weighted ROC-AUC. The most difficult category was still squamous cell carcinoma with 41.67% recall. The framework provided effective multimodal representation and improved class discrimination. However, independent external validation is required to assess its clinical generalizability.

Authors

Institutions

Publication Details

Journal
Electronics
Published
2026-09-20
DOI
https://doi.org/10.3390/electronics15184308
Primary Topic
Cutaneous Melanoma Detection and Management
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Multimodal Swin Transformer with Hyper-Feature and Attention-Gated Clinical Fusion for Skin Lesion Classification

Rahib Hidayat Abiyev, Kamil Dimililer, Rasha Habeeb
Electronics
Cutaneous Melanoma Detection and Management
article

Multimodal Swin Transformer with Hyper-Feature and Attention-Gated Clinical Fusion for Skin Lesion Classification

Rahib Hidayat Abiyev, Kamil Dimililer, Rasha Habeeb
article en

Abstract

Clinical photographs and patient metadata are complementary sources of information for automated skin lesion classification. In this study, a multimodal framework is proposed that combines Swin Transformer Tiny (Swin-T), Hyper-Feature Fusion (HFF), and Attention-Gated Clinical Metadata Fusion (AGCMF). HFF combines four levels of hierarchical representations from the Swin Transformer into a 256-dimensional visual embedding. A multilayer perceptron encodes 20 clinical metadata variables into a 64-dimensional representation. The features are concatenated, adaptively gated and projected for six-class classification. Five-fold patient-grouped cross-validation was performed on PAD-UFES-20, which consists of 2298 clinical photographs of 1373 patients. The proposed model obtained 83.55% accuracy without test-time augmentation (TTA), 83.29% with TTA, 82.92% weighted F1, 77.29% macro F1, 75.75% balanced accuracy, 93.12% macro-ROC-AUC, 94.59% weighted ROC-AUC, and Cohen’s kappa of 0.7702. In the ablation experiments, HFF increased the accuracy of the Swin-T model from 78.64% to 80.91%. Clinical feature concatenation increased it to 81.76%, and the complete HFF + AGCMF framework achieved 83.29%. These results indicate that hierarchical visual fusion and adaptive multimodal integration provide complementary benefits. The proposed framework achieved 3.89 percentage points higher accuracy than the image-plus-metadata baseline under the same TTA evaluation setting. A biopsy-exclusion sensitivity analysis also revealed that the performance of the model was only slightly affected by the removal of the diagnostic-verification variable, with the 19-feature model maintaining 82.91% accuracy, 82.53% weighted F1, and 94.21% weighted ROC-AUC. The most difficult category was still squamous cell carcinoma with 41.67% recall. The framework provided effective multimodal representation and improved class discrimination. However, independent external validation is required to assess its clinical generalizability.

ElectronicsVol. 15(18)
Near East University (CY)
Reduced inequalities
Openalex Percentile: Top 14%
Cutaneous Melanoma Detection and Management
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.