Secure Adaptive Resource Orchestration for Cloud Management with Deep Reinforcement Learning: An Extended Evaluation on Real Traces
Cloud platforms must hold utilisation and latency targets while demand shifts and attack traffic arrive together. Reactive threshold scaling meets neither pressure, and an autoscaler blind to attacks funds the load an adversary requested. Recent work shows adversaries can drive this loop into economic denial of sustainability, so the controller sits inside the attack surface. No prior orchestrator couples workload forecasting, unsupervised anomaly detection and learned scaling in one loop, and none reports multi-seed significance testing. This article extends SARO, presented at IMCOM 2026, to close that gap. We formalise the problem as a Markov decision process with a corrected multi-objective reward, and replace the tabular agent with SARO-DQN, a continuous-state controller trained by three-step Double Q-learning. Across ten held-out days and five seeds, SARO-DQN reaches the highest composite reward (200.1 ± 16.5) and the highest utilisation (68.2%) against eight alternatives, and every reward difference is significant under Welch tests (p<0.05). Ablations attribute 23.6 reward points to the detector (p=1.4×10−6) and 8.9 to the forecast (p=0.016). On UNSW-NB15, the detector attains an AUC of 0.888. Two principles follow. Detectors must consume exogenous traffic-shape signals, and detector and policy must be trained as a coupled system.
Authors
- Priyadarsi Nanda (ORCID: https://orcid.org/0000-0002-5748-155X)
- Hoang Dinh
- Usaid Alibrahem
Institutions
- University of Technology Sydney (AU)
- Najran University (SA)
Publication Details
- Journal
- Electronics
- Published
- 2026-08-31
- DOI
- https://doi.org/10.3390/electronics15173916
- Primary Topic
- Network Security and Intrusion Detection
- Type
- article
- Field-Weighted Citation Impact
- 0.00