The Evolution of Vision Transformers: A Multi‐Dimensional Analysis of Architectural Innovation and Application Domains
ABSTRACT The rapid evolution of deep learning has positioned Vision Transformers (ViTs) as a powerful alternative to convolutional neural networks (CNNs) in computer vision. By leveraging self‐attention to model global dependencies, ViTs achieve state‐of‐the‐art performance across tasks such as image classification and segmentation. This survey presents a comprehensive review of over 150 ViT variants and hybrid models, moving beyond application‐based categorizations to propose a multi‐dimensional taxonomy grounded in architectural evolution and attention design. We critically examine how different mechanisms, such as sparse and linear attention, influence scalability, computational efficiency, and task‐specific accuracy. A dedicated analysis of hybrid CNN‐Transformer architectures reveals how they effectively reintroduced crucial inductive biases, such as locality and translation equivariance, to mitigate the data inefficiency and slow convergence issues inherent in pure ViTs. Furthermore, we extend our review to specialized domains, including 3D analysis, video inpainting, and low‐level vision, correlating architectural choices with performance gains. Here, we show that while ViTs excel at capturing global context, their practical deployment is often constrained by quadratic computational complexity and substantial data requirements. By synthesizing recent breakthroughs and critically evaluating performance trade‐offs, this survey not only provides a structured reference for researchers but also identifies open challenges and delineates promising directions for future research, particularly in developing data‐efficient and resource‐aware Transformer models for real‐world visual computing systems.
Authors
- Piyush Rawat (ORCID: https://orcid.org/0000-0002-6951-7089)
- Divakar Yadav (ORCID: https://orcid.org/0000-0001-6051-479X)
- Prashant Upadhyay (ORCID: https://orcid.org/0000-0002-9257-9181)
- Shobhit Tyagi (ORCID: https://orcid.org/0000-0002-6262-0526)
Institutions
- Jaypee Institute of Information Technology (IN)
- Bennett University (IN)
- Indira Gandhi National Open University (IN)
- Delhi Technological University (IN)
Publication Details
- Journal
- Expert Systems
- Published
- 2026-09-15
- DOI
- https://doi.org/10.1111/exsy.70422
- Primary Topic
- Advanced Neural Network Applications
- Type
- article
- Field-Weighted Citation Impact
- 0.00