Adaptive constrained reinforcement learning for energy-aware scheduling under changing load
Abstract Energy-aware schedulers must trade energy against responsiveness, and reinforcement learning is increasingly used to learn this trade-off. Many such schedulers collapse the two objectives into a single weighted reward. We show that this common choice has an overlooked failure mode: any fixed weight fixes an operating point that drifts as the workload changes, wasting energy at light load and silently breaking latency targets at heavy load. We introduce a scheduler that instead treats the latency-violation rate as an explicit constraint and enforces it with a proportional-integral-derivative controller that continuously retunes a Lagrangian price on violations; the learned policy is additionally conditioned on this price so that a single policy spans many load regimes. Across the tested non-saturated load range, CLASP tracks a five percent violation target closely, and near saturation it degrades gradually rather than collapsing. In a dense sweep of reward weights, no fixed weight maintained the target across all tested loads, and neither does any fixed dual price. CLASP holds the target at substantially lower energy than an always-fast policy, and comes within about ten percent of the cheapest compliant price for its own trained policy, an internal consistency check rather than a bound on achievable efficiency, whereas the evaluated fixed-weight policies drift substantially from the target. The same controller tracks abrupt workload shifts online where open-loop and fixed-price comparators do not, and ablations confirm that both adaptive pricing and price conditioning are necessary. Because the policy is conditioned on the constraint price rather than on a fixed trade-off, one trained policy behaves as a family of controllers: the operator can select any violation target at deployment with no retraining (mean error 0.42 percentage points across targets from two to twenty percent), and the same policy generalizes to loads and bursty arrival patterns never seen in training, holds several heterogeneous service targets at once, and carries over from tabular learning to function approximation. Constraint-driven adaptation thus turns the operating point into a controllable, reusable quantity rather than one baked into a training weight.
Authors
- Khairi Azhar Aziz (ORCID: https://orcid.org/0000-0003-1517-548X)
- Saher Elsayed (ORCID: https://orcid.org/0009-0007-7672-264X)
Institutions
- Universiti Tenaga Nasional (MY)
- University of Pennsylvania (US)
Publication Details
- Journal
- Scientific Reports
- Published
- 2026-09-21
- DOI
- https://doi.org/10.1038/s41598-026-72061-5
- Primary Topic
- Smart Grid Energy Management
- Type
- article
- Field-Weighted Citation Impact
- 0.00