Classification of Brain MRI Images with Vision Transformer: Improving Performance with New Layers and Parameter Optimization

Recent advances in deep learning have revolutionized fields such as robotics, healthcare, and natural language processing. The Vision Transformer (ViT), which operates on self-attention block, has emerged as an alternative to Convolutional Neural Networks. In this study, an improved ViT architecture is proposed for the categorization of brain MRI images as tumorous and non-tumorous. While the standard ViT backbone is utilized for feature encoding, the conventional classification head has been replaced with a multi-stage architecture comprising sequential fully connected layers with 20, 4, and 2 neurons, integrated with Sigmoid and Linear activation functions. This architectural modification, which has been constructed via Greedy Search approach, aims to refine the feature mapping process and enhance the model's sensitivity towards pathological patterns in medical images. To evaluate the performance, a brain MRI dataset has been split into validation, test and training sets, and data augmentation was performed to prevent overfitting. Standard ViT, Swin Transformer, EfficientNet, ResNet and the proposed ViT models have been trained using an ablation technique to optimize network parameters. For the performance analysis, the models have been evaluated using the metrics such as Accuracy, F1-Score, Precision, Recall, AUC values and ROC curves. The results of this work indicate that the proposed ViT’s head structure improves classification success compared to the standard ViT architecture and the other models. Consequently, by redesigning the classification head and optimizing the network, a 13% improvement has been achieved, reaching macro F1-Score of 95.3% for brain image dataset 1 and macro F1-Score of 99.3% for brain image dataset 2.

Authors

Institutions

Publication Details

Journal
Journal of Intelligent Systems Theory and Applications
Published
2026-09-15
DOI
https://doi.org/10.38016/jista.1779498
Primary Topic
Brain Tumor Detection and Classification
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Classification of Brain MRI Images with Vision Transformer: Improving Performance with New Layers and Parameter Optimization

Korhan Günel, İclal Gör, Rıfat AŞLIYAN
Journal of Intelligent Systems Theory and Applications
Brain Tumor Detection and Classification
article

Classification of Brain MRI Images with Vision Transformer: Improving Performance with New Layers and Parameter Optimization

Korhan Günel, İclal Gör, Rıfat AŞLIYAN
article en

Abstract

Recent advances in deep learning have revolutionized fields such as robotics, healthcare, and natural language processing. The Vision Transformer (ViT), which operates on self-attention block, has emerged as an alternative to Convolutional Neural Networks. In this study, an improved ViT architecture is proposed for the categorization of brain MRI images as tumorous and non-tumorous. While the standard ViT backbone is utilized for feature encoding, the conventional classification head has been replaced with a multi-stage architecture comprising sequential fully connected layers with 20, 4, and 2 neurons, integrated with Sigmoid and Linear activation functions. This architectural modification, which has been constructed via Greedy Search approach, aims to refine the feature mapping process and enhance the model's sensitivity towards pathological patterns in medical images. To evaluate the performance, a brain MRI dataset has been split into validation, test and training sets, and data augmentation was performed to prevent overfitting. Standard ViT, Swin Transformer, EfficientNet, ResNet and the proposed ViT models have been trained using an ablation technique to optimize network parameters. For the performance analysis, the models have been evaluated using the metrics such as Accuracy, F1-Score, Precision, Recall, AUC values and ROC curves. The results of this work indicate that the proposed ViT’s head structure improves classification success compared to the standard ViT architecture and the other models. Consequently, by redesigning the classification head and optimizing the network, a 13% improvement has been achieved, reaching macro F1-Score of 95.3% for brain image dataset 1 and macro F1-Score of 99.3% for brain image dataset 2.

Journal of Intelligent Systems Theory and ApplicationsVol. 9(2026)
Adnan Menderes University (TR)
Openalex Percentile: Top 14%
Brain Tumor Detection and Classification
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Classification of Brain MRI Images with Vision Transformer: Improving Performance with New Layers and Parameter Optimization — Korhan Günel, İclal Gör, et al. · Journal of Intelligent Systems Theory and Applications (2026) | TGRS Research Map | TGRS