Permutation-equivariant deep reinforcement learning for resource management and service migration in mobile edge computing
A mobile edge controller has to divide computation and spectrum among users, choose where each task runs, and move user services between sites as devices travel. We treat the three choices as one constrained Markov decision process and solve it with EA-DDPG. The agent is a deep deterministic policy gradient learner whose actor and critic are permutation-equivariant. A single per-user head is shared across users and reads that user’s own features together with pooled site-level and system-level context. Its parameter count therefore does not depend on how many users are present, and the architecture can be trained at any system size without modification. Two other elements separate the agent from an ordinary DDPG application. The action space mixes continuous CPU and bandwidth shares with discrete offloading and migration choices, and we handle it with a deterministic decoder that turns one bounded continuous vector into feasible controls. Per-site capacity constraints are then satisfied automatically and need no penalty term. Service migration is also a decision variable in its own right, and taking it costs the agent a state-transfer delay, an interruption, and transport energy, all of which are charged to the reward. We evaluate on a slotted mobile edge simulator with random-waypoint mobility, Poisson task arrivals, and a cloud tier whose WAN pipe and CPU pool congest as load moves onto them. Against five learning baselines trained under identical conditions, including a conventional flat-MLP DDPG that differs from the proposal only in its encoder, EA-DDPG completes tasks between 78% and 98% faster and misses between 51% and 58% fewer deadlines. Both differences reach the smallest p -value an exact test on ten seeds can produce, \(1.1\times 10^{-5}\) , with complete separation between the two sets of seeds. Against a strong hand-tuned heuristic the outcome is a trade rather than a win: the learned policy uses 53.5% less energy and earns 8.2% more reward, both surviving correction for multiple comparisons, while missing the same fraction of deadlines and taking 11.3% longer on average. That last difference is a disadvantage of the method and it too survives correction. Ablations separate the contribution of each component, including a per-priority-class analysis showing that the priority mechanism halves high-priority completion time while leaving the aggregate violation rate unchanged. A scale sweep from 10 to 100 users separates two senses of scalability: parameter count and per-slot inference stay flat across the range, whereas policy quality degrades under a fixed training budget and the heuristic overtakes the learned policy beyond roughly 40 users. A zero-shot transfer experiment identifies the cause. Because the encoder’s parameters do not depend on the user count, a policy trained at 40 users can be applied unchanged at 100, where it outperforms a policy trained there directly; the limitation is one of sample complexity rather than of the representation. Training diagnostics attribute the flat network’s failure to a collapsed critic, actor gradients an order of magnitude larger, and consequent saturation of 37% of the action coordinates.
Authors
- E. Naresh
- J. Geetha
- Pruthviraj M. Savanur
- Atmuri Sai Mouli
- Ba. Ba. Ibrahim
- Benadin Benny
Institutions
- M S Ramaiah University of Applied Sciences (IN)
- Ramaiah Institute of Technology (IN)
Publication Details
- Journal
- Scientific Reports
- Published
- 2026-09-24
- DOI
- https://doi.org/10.1038/s41598-026-72506-x
- Primary Topic
- IoT and Edge/Fog Computing
- Type
- article
- Field-Weighted Citation Impact
- 0.00