A controlled benchmark of CNN architectures for MRI-based brain tumor detection: Custom networks and transfer learning versus Vision Transformer baselines under a unified image-processing pipeline

Magnetic resonance imaging (MRI) is the reference modality for non-invasive brain-tumor assessment, but visual interpretation is slow, subject to inter-observer variability and hard to scale. Computer-aided systems based on convolutional neural networks (CNNs) can help, yet the biomedical image-processing literature lacks fair, head-to-head comparisons of custom and pre-trained models under a single pipeline. We present a controlled benchmark of twelve architectures — six custom CNNs, four pre-trained transfer-learning CNNs (VGG-16, VGG-19, ResNet-50 and MobileNetV2) and two Vision Transformers (ViT-B/16 and Swin-Tiny) — for binary tumor-versus-no-tumor classification in MRI, all trained under the same data, augmentation, splits, optimizer and learning-rate sweep, with model-appropriate input normalization. The Vision Transformer ViT-B/16 attains the best accuracy (0.9665) and F1 score (0.9662), whereas the hierarchical Swin-Tiny reaches only 0.9297 and does not beat the CNNs, indicating that the transformer advantage is architecture-specific rather than generic. With backbone-appropriate preprocessing the transfer-learning CNNs are strong (VGG-16 0.9521, ResNet-50 0.9473) and the best custom network trained from scratch (Arch-3, 0.9489) is competitive, so the gap between families is narrow; MobileNetV2 offers the best accuracy–efficiency trade-off. A preprocessing ablation shows that input normalization is a decisive, commonly neglected factor for frozen transfer learning: under a naive [ 0 , 1 ] pipeline ResNet-50 collapses to 0.8147 and Swin-Tiny to 0.8546, while MobileNetV2 and ViT-B/16 remain robust. Grad-CAM confirms that MobileNetV2 attends to tumor regions. The benchmark, released with public code, offers reproducible guidance for selecting architectures in clinical decision-support and edge-deployment scenarios.

Authors

Institutions

Publication Details

Journal
Biomedical Signal Processing and Control
Published
2026-09-14
DOI
https://doi.org/10.1016/j.bspc.2026.111408
Primary Topic
Brain Tumor Detection and Classification
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

A controlled benchmark of CNN architectures for MRI-based brain tumor detection: Custom networks and transfer learning versus Vision Transformer baselines under a unified image-processing pipeline

Ricardo S. Alonso, Emilio Yoshihiro Rosanes Nishiyama
Biomedical Signal Processing and Control
Brain Tumor Detection and Classification
article

A controlled benchmark of CNN architectures for MRI-based brain tumor detection: Custom networks and transfer learning versus Vision Transformer baselines under a unified image-processing pipeline

Ricardo S. Alonso, Emilio Yoshihiro Rosanes Nishiyama
article en

Abstract

Magnetic resonance imaging (MRI) is the reference modality for non-invasive brain-tumor assessment, but visual interpretation is slow, subject to inter-observer variability and hard to scale. Computer-aided systems based on convolutional neural networks (CNNs) can help, yet the biomedical image-processing literature lacks fair, head-to-head comparisons of custom and pre-trained models under a single pipeline. We present a controlled benchmark of twelve architectures — six custom CNNs, four pre-trained transfer-learning CNNs (VGG-16, VGG-19, ResNet-50 and MobileNetV2) and two Vision Transformers (ViT-B/16 and Swin-Tiny) — for binary tumor-versus-no-tumor classification in MRI, all trained under the same data, augmentation, splits, optimizer and learning-rate sweep, with model-appropriate input normalization. The Vision Transformer ViT-B/16 attains the best accuracy (0.9665) and F1 score (0.9662), whereas the hierarchical Swin-Tiny reaches only 0.9297 and does not beat the CNNs, indicating that the transformer advantage is architecture-specific rather than generic. With backbone-appropriate preprocessing the transfer-learning CNNs are strong (VGG-16 0.9521, ResNet-50 0.9473) and the best custom network trained from scratch (Arch-3, 0.9489) is competitive, so the gap between families is narrow; MobileNetV2 offers the best accuracy–efficiency trade-off. A preprocessing ablation shows that input normalization is a decisive, commonly neglected factor for frozen transfer learning: under a naive [ 0 , 1 ] pipeline ResNet-50 collapses to 0.8147 and Swin-Tiny to 0.8546, while MobileNetV2 and ViT-B/16 remain robust. Grad-CAM confirms that MobileNetV2 attends to tumor regions. The benchmark, released with public code, offers reproducible guidance for selecting architectures in clinical decision-support and edge-deployment scenarios.

Biomedical Signal Processing and ControlVol. 129
Fundación INTRAS (ES), Universidad Internacional De La Rioja (ES), Universidad Europea de Madrid (ES)
Peace, Justice and strong institutions
Openalex Percentile: Top 13%
Brain Tumor Detection and Classification
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.