A medical image classification algorithm based on a hierarchical and complementary attention-enhanced Swin Transformer model

Abstract With the rapid development of precision medicine and intelligent diagnostic technologies, automatic medical image classification has become an important tool for assisting clinical decision-making. However, substantial variations in lesion scale, complex long-range dependencies of tissue structures, and the need to capture subtle anatomical features present significant challenges to existing deep learning models. To address the limitations of the Swin Transformer in modeling multi-scale lesions and multi-level feature interactions, this study proposes a hierarchical complementary feature enhancement framework based on the Swin Transformer for medical image classification. The proposed architecture performs collaborative feature learning at three representation levels, including macro-scale lesion perception, global contextual interaction, and local detail refinement. Specifically, a Multi-scale Depthwise SE Block (MSD-SE Block) is introduced at the input of each stage of the Swin Transformer to enhance the model’s multi-scale feature representation capability. Subsequently, a Residual Convolutional Attention (RCA) module is integrated following the self-attention mechanism and the Multi-Layer Perceptron (MLP) to strengthen global contextual modeling, while a Local Detail Enhanced Residual Channel-Spatial Attention (LDERCSA) module is employed to refine subtle anatomical structures and discriminative local features. Through the coordinated interaction of these components, the proposed framework establishes a hierarchical feature enhancement mechanism that effectively improves medical image representation across multiple scales and feature levels. Comprehensive experiments were conducted on eight core subsets of MedMNIST v2, including BloodMNIST, BreastMNIST, DermaMNIST, OCTMNIST, OrganSMNIST, PathMNIST, PneumoniaMNIST, and RetinaMNIST, using an input resolution of 224 $$\\times$$ 224. Single-module comparison and ablation studies demonstrate that each proposed component contributes positively to the overall performance. Experimental results show that the proposed model achieves significant performance improvements on BreastMNIST, OCTMNIST, OrganSMNIST, and PneumoniaMNIST. Furthermore, comparative evaluations across all datasets confirm that the proposed framework exhibits strong robustness and generalization capability when processing multiple medical imaging modalities, including ultrasound, CT, X-ray, endoscopic, and microscopic images, thereby providing reliable technical support for computer-aided medical diagnosis systems.

Authors

Institutions

Publication Details

Journal
Scientific Reports
Published
2026-08-31
DOI
https://doi.org/10.1038/s41598-026-69149-3
Primary Topic
AI in cancer detection
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

A medical image classification algorithm based on a hierarchical and complementary attention-enhanced Swin Transformer model

Mingzhan Zhao, Yi Zhang, Yachao Si
Scientific Reports
AI in cancer detection
article

A medical image classification algorithm based on a hierarchical and complementary attention-enhanced Swin Transformer model

Mingzhan Zhao, Yi Zhang, Yachao Si
article en

Abstract

Abstract With the rapid development of precision medicine and intelligent diagnostic technologies, automatic medical image classification has become an important tool for assisting clinical decision-making. However, substantial variations in lesion scale, complex long-range dependencies of tissue structures, and the need to capture subtle anatomical features present significant challenges to existing deep learning models. To address the limitations of the Swin Transformer in modeling multi-scale lesions and multi-level feature interactions, this study proposes a hierarchical complementary feature enhancement framework based on the Swin Transformer for medical image classification. The proposed architecture performs collaborative feature learning at three representation levels, including macro-scale lesion perception, global contextual interaction, and local detail refinement. Specifically, a Multi-scale Depthwise SE Block (MSD-SE Block) is introduced at the input of each stage of the Swin Transformer to enhance the model’s multi-scale feature representation capability. Subsequently, a Residual Convolutional Attention (RCA) module is integrated following the self-attention mechanism and the Multi-Layer Perceptron (MLP) to strengthen global contextual modeling, while a Local Detail Enhanced Residual Channel-Spatial Attention (LDERCSA) module is employed to refine subtle anatomical structures and discriminative local features. Through the coordinated interaction of these components, the proposed framework establishes a hierarchical feature enhancement mechanism that effectively improves medical image representation across multiple scales and feature levels. Comprehensive experiments were conducted on eight core subsets of MedMNIST v2, including BloodMNIST, BreastMNIST, DermaMNIST, OCTMNIST, OrganSMNIST, PathMNIST, PneumoniaMNIST, and RetinaMNIST, using an input resolution of 224 $$\times$$ 224. Single-module comparison and ablation studies demonstrate that each proposed component contributes positively to the overall performance. Experimental results show that the proposed model achieves significant performance improvements on BreastMNIST, OCTMNIST, OrganSMNIST, and PneumoniaMNIST. Furthermore, comparative evaluations across all datasets confirm that the proposed framework exhibits strong robustness and generalization capability when processing multiple medical imaging modalities, including ultrasound, CT, X-ray, endoscopic, and microscopic images, thereby providing reliable technical support for computer-aided medical diagnosis systems.

Scientific Reports
Zhangjiakou Academy of Agricultural Sciences (CN), Hebei University of Architecture (CN)
Reduced inequalities
Openalex Percentile: Top 8%
AI in cancer detection
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.