Review and optimization of image classification models for edge AI applications
Image classification on edge devices requires models that maintain accuracy while reducing inference cost. This paper reviews the development of image classification models from convolutional neural networks to Vision Transformer models and examines their relevance to edge artificial intelligence applications. To improve Vision Transformer inference efficiency without redesigning the backbone, we propose a selective token merging method that identifies low-significance image tokens and applies token merging mainly to those tokens. Experimental results on CIFAR-10 using ViT-Tiny show that the proposed method improves the accuracy-throughput tradeoff compared with standard Token Merging. In particular, selectively merging the least significant token group with a layer-dependent schedule achieves 97.41% accuracy and 19.04 images/s, compared with 97.22% accuracy and 18.97 images/s for standard Token Merging. Additional ImageNet-1 K evaluations using ViT-Tiny, ViT-Small, and ViT-Base show that selective merging in an shallow-layer window preserves more of the baseline accuracy than standard ToMe while retaining an inference speedup. These results show that token significance can guide more accurate and efficient token reduction for edge-oriented image classification.
Authors
- Sungjae An (ORCID: https://orcid.org/0000-0003-2033-6049)
- Sungju Ryu (ORCID: https://orcid.org/0000-0002-0254-391X)
- Sunghyun Kim
- Huseok Lee
- Hanbin Cho
- Taejee Kim
Institutions
- Sogang University (KR)
Publication Details
- Journal
- Nano Convergence
- Published
- 2026-09-22
- DOI
- https://doi.org/10.1186/s40580-026-00574-w
- Primary Topic
- Advanced Neural Network Applications
- Type
- article
- Field-Weighted Citation Impact
- 0.00