Review and optimization of image classification models for edge AI applications

Image classification on edge devices requires models that maintain accuracy while reducing inference cost. This paper reviews the development of image classification models from convolutional neural networks to Vision Transformer models and examines their relevance to edge artificial intelligence applications. To improve Vision Transformer inference efficiency without redesigning the backbone, we propose a selective token merging method that identifies low-significance image tokens and applies token merging mainly to those tokens. Experimental results on CIFAR-10 using ViT-Tiny show that the proposed method improves the accuracy-throughput tradeoff compared with standard Token Merging. In particular, selectively merging the least significant token group with a layer-dependent schedule achieves 97.41% accuracy and 19.04 images/s, compared with 97.22% accuracy and 18.97 images/s for standard Token Merging. Additional ImageNet-1 K evaluations using ViT-Tiny, ViT-Small, and ViT-Base show that selective merging in an shallow-layer window preserves more of the baseline accuracy than standard ToMe while retaining an inference speedup. These results show that token significance can guide more accurate and efficient token reduction for edge-oriented image classification.

Authors

Institutions

Publication Details

Journal
Nano Convergence
Published
2026-09-22
DOI
https://doi.org/10.1186/s40580-026-00574-w
Primary Topic
Advanced Neural Network Applications
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Review and optimization of image classification models for edge AI applications

Sungjae An, Sungju Ryu, Sunghyun Kim, Huseok Lee et al.
Nano Convergence
Advanced Neural Network Applications
article

Review and optimization of image classification models for edge AI applications

Sungjae An, Sungju Ryu, Sunghyun Kim, Huseok Lee, Hanbin Cho, Taejee Kim
article en

Abstract

Image classification on edge devices requires models that maintain accuracy while reducing inference cost. This paper reviews the development of image classification models from convolutional neural networks to Vision Transformer models and examines their relevance to edge artificial intelligence applications. To improve Vision Transformer inference efficiency without redesigning the backbone, we propose a selective token merging method that identifies low-significance image tokens and applies token merging mainly to those tokens. Experimental results on CIFAR-10 using ViT-Tiny show that the proposed method improves the accuracy-throughput tradeoff compared with standard Token Merging. In particular, selectively merging the least significant token group with a layer-dependent schedule achieves 97.41% accuracy and 19.04 images/s, compared with 97.22% accuracy and 18.97 images/s for standard Token Merging. Additional ImageNet-1 K evaluations using ViT-Tiny, ViT-Small, and ViT-Base show that selective merging in an shallow-layer window preserves more of the baseline accuracy than standard ToMe while retaining an inference speedup. These results show that token significance can guide more accurate and efficient token reduction for edge-oriented image classification.

Nano ConvergenceVol. 13(1)
Sogang University (KR)
Openalex Percentile: Top 13%
Advanced Neural Network Applications
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.