ProfEdge: Efficient Construction of DNN Performance Evaluation Model on Edge Devices

With the rapid development of the Internet of Things (IoT) and artificial intelligence (AI) technologies, edge computing emerges as a crucial computing paradigm. By processing data near the source, edge computing enables faster and more efficient intelligent services. However, edge devices usually limit computational resources, and existing DNN inference latency profiling methods often rely on internal model details or large-scale latency measurements, making them costly and unsuitable for black-box deployment scenarios. This paper proposes ProfEdge, a fast construction framework for deep neural network (DNN) latency profiling models based on Gaussian process regression and Bayesian optimization. ProfEdge adaptively samples real latency measurements under different batch-size states and tunes profiling-model hyperparameters to reduce construction cost while improving profiling accuracy. Specifically, ProfEdge builds an adaptive sampling module based on Gaussian process regression to locate high-error regions through coarse-grained sampling and dynamically refine the sampling process. It further designs a dynamic Bayesian optimization mechanism to improve the accuracy of the latency profiling model. Finally, ProfEdge constructs a cross-device performance mapping model to migrate an existing profiling model to a target device with lightweight stratified calibration, thereby avoiding full reconstruction of the target-device profiling model. Experiments on various edge devices and DNN models, including CNN-based and transformer-based workloads, show that ProfEdge reduces profiling errors by up to 80% and saves over 70% of profiling construction cost compared with existing methods. The cross-device migration results further demonstrate that ProfEdge can achieve competitive profiling accuracy on new devices with only a small number of target-device calibration samples.

Authors

Publication Details

Journal
ACM Transactions on Internet Technology
Published
2026-09-11
DOI
https://doi.org/10.1145/3847112
Primary Topic
IoT and Edge/Fog Computing
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

ProfEdge: Efficient Construction of DNN Performance Evaluation Model on Edge Devices

Borui Li, Songtao Lu, Shuai Wang, Weilong Wang et al.
ACM Transactions on Internet Technology
IoT and Edge/Fog Computing
article

ProfEdge: Efficient Construction of DNN Performance Evaluation Model on Edge Devices

Borui Li, Songtao Lu, Shuai Wang, Weilong Wang, Zhao‐Dong Xu, Jiawei Liu
article en

Abstract

With the rapid development of the Internet of Things (IoT) and artificial intelligence (AI) technologies, edge computing emerges as a crucial computing paradigm. By processing data near the source, edge computing enables faster and more efficient intelligent services. However, edge devices usually limit computational resources, and existing DNN inference latency profiling methods often rely on internal model details or large-scale latency measurements, making them costly and unsuitable for black-box deployment scenarios. This paper proposes ProfEdge, a fast construction framework for deep neural network (DNN) latency profiling models based on Gaussian process regression and Bayesian optimization. ProfEdge adaptively samples real latency measurements under different batch-size states and tunes profiling-model hyperparameters to reduce construction cost while improving profiling accuracy. Specifically, ProfEdge builds an adaptive sampling module based on Gaussian process regression to locate high-error regions through coarse-grained sampling and dynamically refine the sampling process. It further designs a dynamic Bayesian optimization mechanism to improve the accuracy of the latency profiling model. Finally, ProfEdge constructs a cross-device performance mapping model to migrate an existing profiling model to a target device with lightweight stratified calibration, thereby avoiding full reconstruction of the target-device profiling model. Experiments on various edge devices and DNN models, including CNN-based and transformer-based workloads, show that ProfEdge reduces profiling errors by up to 80% and saves over 70% of profiling construction cost compared with existing methods. The cross-device migration results further demonstrate that ProfEdge can achieve competitive profiling accuracy on new devices with only a small number of target-device calibration samples.

ACM Transactions on Internet Technology
Openalex Percentile: Top 8%
IoT and Edge/Fog Computing
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.