The Evolution of Vision Transformers: A Multi‐Dimensional Analysis of Architectural Innovation and Application Domains

ABSTRACT The rapid evolution of deep learning has positioned Vision Transformers (ViTs) as a powerful alternative to convolutional neural networks (CNNs) in computer vision. By leveraging self‐attention to model global dependencies, ViTs achieve state‐of‐the‐art performance across tasks such as image classification and segmentation. This survey presents a comprehensive review of over 150 ViT variants and hybrid models, moving beyond application‐based categorizations to propose a multi‐dimensional taxonomy grounded in architectural evolution and attention design. We critically examine how different mechanisms, such as sparse and linear attention, influence scalability, computational efficiency, and task‐specific accuracy. A dedicated analysis of hybrid CNN‐Transformer architectures reveals how they effectively reintroduced crucial inductive biases, such as locality and translation equivariance, to mitigate the data inefficiency and slow convergence issues inherent in pure ViTs. Furthermore, we extend our review to specialized domains, including 3D analysis, video inpainting, and low‐level vision, correlating architectural choices with performance gains. Here, we show that while ViTs excel at capturing global context, their practical deployment is often constrained by quadratic computational complexity and substantial data requirements. By synthesizing recent breakthroughs and critically evaluating performance trade‐offs, this survey not only provides a structured reference for researchers but also identifies open challenges and delineates promising directions for future research, particularly in developing data‐efficient and resource‐aware Transformer models for real‐world visual computing systems.

Authors

Institutions

Publication Details

Journal
Expert Systems
Published
2026-09-15
DOI
https://doi.org/10.1111/exsy.70422
Primary Topic
Advanced Neural Network Applications
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

The Evolution of Vision Transformers: A Multi‐Dimensional Analysis of Architectural Innovation and Application Domains

Piyush Rawat, Divakar Yadav, Prashant Upadhyay, Shobhit Tyagi
Expert Systems
Advanced Neural Network Applications
article

The Evolution of Vision Transformers: A Multi‐Dimensional Analysis of Architectural Innovation and Application Domains

Piyush Rawat, Divakar Yadav, Prashant Upadhyay, Shobhit Tyagi
article en

Abstract

ABSTRACT The rapid evolution of deep learning has positioned Vision Transformers (ViTs) as a powerful alternative to convolutional neural networks (CNNs) in computer vision. By leveraging self‐attention to model global dependencies, ViTs achieve state‐of‐the‐art performance across tasks such as image classification and segmentation. This survey presents a comprehensive review of over 150 ViT variants and hybrid models, moving beyond application‐based categorizations to propose a multi‐dimensional taxonomy grounded in architectural evolution and attention design. We critically examine how different mechanisms, such as sparse and linear attention, influence scalability, computational efficiency, and task‐specific accuracy. A dedicated analysis of hybrid CNN‐Transformer architectures reveals how they effectively reintroduced crucial inductive biases, such as locality and translation equivariance, to mitigate the data inefficiency and slow convergence issues inherent in pure ViTs. Furthermore, we extend our review to specialized domains, including 3D analysis, video inpainting, and low‐level vision, correlating architectural choices with performance gains. Here, we show that while ViTs excel at capturing global context, their practical deployment is often constrained by quadratic computational complexity and substantial data requirements. By synthesizing recent breakthroughs and critically evaluating performance trade‐offs, this survey not only provides a structured reference for researchers but also identifies open challenges and delineates promising directions for future research, particularly in developing data‐efficient and resource‐aware Transformer models for real‐world visual computing systems.

Expert SystemsVol. 43(10)
Jaypee Institute of Information Technology (IN), Bennett University (IN), Indira Gandhi National Open University (IN), Delhi Technological University (IN)
Industry, innovation and infrastructure
Openalex Percentile: Top 13%
Advanced Neural Network Applications
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.