Fine-GrainedEnergy Assessment of Vision Transformer Operators on Hybrid SoC Platforms: A CPU, iGPU, and NPU Case Study

As CPUs alone are often insufficiently efficient for running Vision Transformers (ViTs), contemporary computing platforms, particularly edge systems, increasingly adopt heterogeneous architectures that combine CPUs with GPUs and Neural Processing Units (NPUs). This paper presents a detailed energy assessment of individual ViT functions running on a modern hybrid processor, the AMD Ryzen AI 9 HX 370, which features a CPU, an integrated GPU (iGPU), and a dedicated Neural Processing Unit (NPU). Unlike macro-level benchmarks, our analysis takes a fine-grained approach by focusing on the most widely used ViT functions. We measure the execution time, average power draw, and idle-subtracted energy of key mathematical and structural layers such as matrix multiplications, attention mechanisms, normalization, and activations on each engine, and we profile the CPU a second time under matched INT8 quantization so that architectural effects can be separated from precision effects. Our findings reveal clear, hardware-specific trade-offs. The NPU is substantially more efficient on dense matrix multiplication (4.5× on the feed-forward projection, about 2.9× on the attention projections) and on strided convolution (up to 9.5×), and the advantage persists against the CPU at the CPU’s own cheaper precision, so it is not attributable to the NPU’s lower precision alone. The CPU remains more efficient for elementwise, activation, and normalization operators, in some cases by an order of magnitude. Two results run against common expectation: quantization was a net energy cost on fixed hardware for thirteen of sixteen isolated operators and for the fused attention block, because quantize and dequantize overhead dominates where arithmetic per element is low; and at long context, softmax consumed 26% of the attention block’s energy while performing 1.3% of its arithmetic, a divergence no operation-count model recovers. These insights provide concrete guidelines for developers designing energy-aware runtime schedulers, showing exactly when to offload specific transformer operators to maximize efficiency on hybrid client hardware.

Authors

Institutions

Publication Details

Journal
Applied Sciences
Published
2026-09-28
DOI
https://doi.org/10.3390/app16199626
Primary Topic
Advanced Neural Network Applications
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Fine-GrainedEnergy Assessment of Vision Transformer Operators on Hybrid SoC Platforms: A CPU, iGPU, and NPU Case Study

Christoforos Kachris, Lincoln Ho
Applied Sciences
Advanced Neural Network Applications
article

Fine-GrainedEnergy Assessment of Vision Transformer Operators on Hybrid SoC Platforms: A CPU, iGPU, and NPU Case Study

Christoforos Kachris, Lincoln Ho
article en

Abstract

As CPUs alone are often insufficiently efficient for running Vision Transformers (ViTs), contemporary computing platforms, particularly edge systems, increasingly adopt heterogeneous architectures that combine CPUs with GPUs and Neural Processing Units (NPUs). This paper presents a detailed energy assessment of individual ViT functions running on a modern hybrid processor, the AMD Ryzen AI 9 HX 370, which features a CPU, an integrated GPU (iGPU), and a dedicated Neural Processing Unit (NPU). Unlike macro-level benchmarks, our analysis takes a fine-grained approach by focusing on the most widely used ViT functions. We measure the execution time, average power draw, and idle-subtracted energy of key mathematical and structural layers such as matrix multiplications, attention mechanisms, normalization, and activations on each engine, and we profile the CPU a second time under matched INT8 quantization so that architectural effects can be separated from precision effects. Our findings reveal clear, hardware-specific trade-offs. The NPU is substantially more efficient on dense matrix multiplication (4.5× on the feed-forward projection, about 2.9× on the attention projections) and on strided convolution (up to 9.5×), and the advantage persists against the CPU at the CPU’s own cheaper precision, so it is not attributable to the NPU’s lower precision alone. The CPU remains more efficient for elementwise, activation, and normalization operators, in some cases by an order of magnitude. Two results run against common expectation: quantization was a net energy cost on fixed hardware for thirteen of sixteen isolated operators and for the fused attention block, because quantize and dequantize overhead dominates where arithmetic per element is low; and at long context, softmax consumed 26% of the attention block’s energy while performing 1.3% of its arithmetic, a divergence no operation-count model recovers. These insights provide concrete guidelines for developers designing energy-aware runtime schedulers, showing exactly when to offload specific transformer operators to maximize efficiency on hybrid client hardware.

Applied SciencesVol. 16(19)
Princeton University (US), University of West Attica (GR)
Affordable and clean energy
Openalex Percentile: Top 14%
Advanced Neural Network Applications
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.