Permutation-equivariant deep reinforcement learning for resource management and service migration in mobile edge computing

A mobile edge controller has to divide computation and spectrum among users, choose where each task runs, and move user services between sites as devices travel. We treat the three choices as one constrained Markov decision process and solve it with EA-DDPG. The agent is a deep deterministic policy gradient learner whose actor and critic are permutation-equivariant. A single per-user head is shared across users and reads that user’s own features together with pooled site-level and system-level context. Its parameter count therefore does not depend on how many users are present, and the architecture can be trained at any system size without modification. Two other elements separate the agent from an ordinary DDPG application. The action space mixes continuous CPU and bandwidth shares with discrete offloading and migration choices, and we handle it with a deterministic decoder that turns one bounded continuous vector into feasible controls. Per-site capacity constraints are then satisfied automatically and need no penalty term. Service migration is also a decision variable in its own right, and taking it costs the agent a state-transfer delay, an interruption, and transport energy, all of which are charged to the reward. We evaluate on a slotted mobile edge simulator with random-waypoint mobility, Poisson task arrivals, and a cloud tier whose WAN pipe and CPU pool congest as load moves onto them. Against five learning baselines trained under identical conditions, including a conventional flat-MLP DDPG that differs from the proposal only in its encoder, EA-DDPG completes tasks between 78% and 98% faster and misses between 51% and 58% fewer deadlines. Both differences reach the smallest p -value an exact test on ten seeds can produce, \(1.1\times 10^{-5}\) , with complete separation between the two sets of seeds. Against a strong hand-tuned heuristic the outcome is a trade rather than a win: the learned policy uses 53.5% less energy and earns 8.2% more reward, both surviving correction for multiple comparisons, while missing the same fraction of deadlines and taking 11.3% longer on average. That last difference is a disadvantage of the method and it too survives correction. Ablations separate the contribution of each component, including a per-priority-class analysis showing that the priority mechanism halves high-priority completion time while leaving the aggregate violation rate unchanged. A scale sweep from 10 to 100 users separates two senses of scalability: parameter count and per-slot inference stay flat across the range, whereas policy quality degrades under a fixed training budget and the heuristic overtakes the learned policy beyond roughly 40 users. A zero-shot transfer experiment identifies the cause. Because the encoder’s parameters do not depend on the user count, a policy trained at 40 users can be applied unchanged at 100, where it outperforms a policy trained there directly; the limitation is one of sample complexity rather than of the representation. Training diagnostics attribute the flat network’s failure to a collapsed critic, actor gradients an order of magnitude larger, and consequent saturation of 37% of the action coordinates.

Authors

Institutions

Publication Details

Journal
Scientific Reports
Published
2026-09-24
DOI
https://doi.org/10.1038/s41598-026-72506-x
Primary Topic
IoT and Edge/Fog Computing
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Permutation-equivariant deep reinforcement learning for resource management and service migration in mobile edge computing

E. Naresh, J. Geetha, Pruthviraj M. Savanur, Atmuri Sai Mouli et al.
Scientific Reports
IoT and Edge/Fog Computing
article

Permutation-equivariant deep reinforcement learning for resource management and service migration in mobile edge computing

E. Naresh, J. Geetha, Pruthviraj M. Savanur, Atmuri Sai Mouli, Ba. Ba. Ibrahim, Benadin Benny
article en

Abstract

A mobile edge controller has to divide computation and spectrum among users, choose where each task runs, and move user services between sites as devices travel. We treat the three choices as one constrained Markov decision process and solve it with EA-DDPG. The agent is a deep deterministic policy gradient learner whose actor and critic are permutation-equivariant. A single per-user head is shared across users and reads that user’s own features together with pooled site-level and system-level context. Its parameter count therefore does not depend on how many users are present, and the architecture can be trained at any system size without modification. Two other elements separate the agent from an ordinary DDPG application. The action space mixes continuous CPU and bandwidth shares with discrete offloading and migration choices, and we handle it with a deterministic decoder that turns one bounded continuous vector into feasible controls. Per-site capacity constraints are then satisfied automatically and need no penalty term. Service migration is also a decision variable in its own right, and taking it costs the agent a state-transfer delay, an interruption, and transport energy, all of which are charged to the reward. We evaluate on a slotted mobile edge simulator with random-waypoint mobility, Poisson task arrivals, and a cloud tier whose WAN pipe and CPU pool congest as load moves onto them. Against five learning baselines trained under identical conditions, including a conventional flat-MLP DDPG that differs from the proposal only in its encoder, EA-DDPG completes tasks between 78% and 98% faster and misses between 51% and 58% fewer deadlines. Both differences reach the smallest p -value an exact test on ten seeds can produce, \(1.1\times 10^{-5}\) , with complete separation between the two sets of seeds. Against a strong hand-tuned heuristic the outcome is a trade rather than a win: the learned policy uses 53.5% less energy and earns 8.2% more reward, both surviving correction for multiple comparisons, while missing the same fraction of deadlines and taking 11.3% longer on average. That last difference is a disadvantage of the method and it too survives correction. Ablations separate the contribution of each component, including a per-priority-class analysis showing that the priority mechanism halves high-priority completion time while leaving the aggregate violation rate unchanged. A scale sweep from 10 to 100 users separates two senses of scalability: parameter count and per-slot inference stay flat across the range, whereas policy quality degrades under a fixed training budget and the heuristic overtakes the learned policy beyond roughly 40 users. A zero-shot transfer experiment identifies the cause. Because the encoder’s parameters do not depend on the user count, a policy trained at 40 users can be applied unchanged at 100, where it outperforms a policy trained there directly; the limitation is one of sample complexity rather than of the representation. Training diagnostics attribute the flat network’s failure to a collapsed critic, actor gradients an order of magnitude larger, and consequent saturation of 37% of the action coordinates.

Scientific Reports
M S Ramaiah University of Applied Sciences (IN), Ramaiah Institute of Technology (IN)
Reduced inequalities
Openalex Percentile: Top 9%
IoT and Edge/Fog Computing
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.