Secure Adaptive Resource Orchestration for Cloud Management with Deep Reinforcement Learning: An Extended Evaluation on Real Traces

Cloud platforms must hold utilisation and latency targets while demand shifts and attack traffic arrive together. Reactive threshold scaling meets neither pressure, and an autoscaler blind to attacks funds the load an adversary requested. Recent work shows adversaries can drive this loop into economic denial of sustainability, so the controller sits inside the attack surface. No prior orchestrator couples workload forecasting, unsupervised anomaly detection and learned scaling in one loop, and none reports multi-seed significance testing. This article extends SARO, presented at IMCOM 2026, to close that gap. We formalise the problem as a Markov decision process with a corrected multi-objective reward, and replace the tabular agent with SARO-DQN, a continuous-state controller trained by three-step Double Q-learning. Across ten held-out days and five seeds, SARO-DQN reaches the highest composite reward (200.1 ± 16.5) and the highest utilisation (68.2%) against eight alternatives, and every reward difference is significant under Welch tests (p<0.05). Ablations attribute 23.6 reward points to the detector (p=1.4×10−6) and 8.9 to the forecast (p=0.016). On UNSW-NB15, the detector attains an AUC of 0.888. Two principles follow. Detectors must consume exogenous traffic-shape signals, and detector and policy must be trained as a coupled system.

Authors

Institutions

Publication Details

Journal
Electronics
Published
2026-08-31
DOI
https://doi.org/10.3390/electronics15173916
Primary Topic
Network Security and Intrusion Detection
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Secure Adaptive Resource Orchestration for Cloud Management with Deep Reinforcement Learning: An Extended Evaluation on Real Traces

Priyadarsi Nanda, Hoang Dinh, Usaid Alibrahem
Electronics
Network Security and Intrusion Detection
article

Secure Adaptive Resource Orchestration for Cloud Management with Deep Reinforcement Learning: An Extended Evaluation on Real Traces

Priyadarsi Nanda, Hoang Dinh, Usaid Alibrahem
article en

Abstract

Cloud platforms must hold utilisation and latency targets while demand shifts and attack traffic arrive together. Reactive threshold scaling meets neither pressure, and an autoscaler blind to attacks funds the load an adversary requested. Recent work shows adversaries can drive this loop into economic denial of sustainability, so the controller sits inside the attack surface. No prior orchestrator couples workload forecasting, unsupervised anomaly detection and learned scaling in one loop, and none reports multi-seed significance testing. This article extends SARO, presented at IMCOM 2026, to close that gap. We formalise the problem as a Markov decision process with a corrected multi-objective reward, and replace the tabular agent with SARO-DQN, a continuous-state controller trained by three-step Double Q-learning. Across ten held-out days and five seeds, SARO-DQN reaches the highest composite reward (200.1 ± 16.5) and the highest utilisation (68.2%) against eight alternatives, and every reward difference is significant under Welch tests (p<0.05). Ablations attribute 23.6 reward points to the detector (p=1.4×10−6) and 8.9 to the forecast (p=0.016). On UNSW-NB15, the detector attains an AUC of 0.888. Two principles follow. Detectors must consume exogenous traffic-shape signals, and detector and policy must be trained as a coupled system.

ElectronicsVol. 15(17)
University of Technology Sydney (AU), Najran University (SA)
Responsible consumption and production
Openalex Percentile: Top 8%
Network Security and Intrusion Detection
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Secure Adaptive Resource Orchestration for Cloud Management with Deep Reinforcement Learning: An Extended Evaluation on Real Traces — Priyadarsi Nanda, Hoang Dinh, et al. · Electronics (2026) | TGRS Research Map | TGRS