Uncertainty-aware hybrid autoscaling with bias-corrected deep forecasting for cost-efficient capacity-shortfall control across cloud workload traces

We present HybridTuned, an uncertainty-aware hybrid autoscaling framework designed to mitigate bias-induced capacity shortfalls when continuous workload forecasts are translated into discrete provisioning decisions. The primary evaluation uses the Google 2019 Cluster sample, aggregated at 5-minute intervals and partitioned chronologically into training, validation, and held-out test blocks. Point-forecasting performance is compared across LSTM, vanilla Transformer, and PatchTST models over five training seeds using Holm-adjusted Diebold–Mariano tests, while a quantile LSTM baseline is assessed separately using pinball loss and empirical coverage. HybridTuned combines median-based forecast-bias correction, a scale-normalized uncertainty margin derived from positive bias-corrected validation residuals, and a utilization-triggered reactive override. Controller parameters are selected exclusively through validation replay under an explicit capacity-shortfall feasibility constraint, and uncertainty in controller outcomes is quantified using paired moving-block bootstrap analysis. On the 1353-interval Google test trajectory, HybridTuned reduces replica-equivalent machine-hours from 5200.83 to 4,638.58, corresponding to a 10.81% reduction relative to the Reactive baseline. Both policies produce one shortage interval, yielding the same under-provisioning percentage (UnderPct) of 0.07391%. HybridTuned also reduces mean over-provisioning from 19.806 to 14.819 and switching frequency from 3.1308 to 0.2927 switches per hour. Cross-dataset evaluation on the GWA/Bitbrains traces yields an 8.0% cost reduction, decreases mean waste from 17.42 to 13.96, and reduces switching frequency from 2.41 to 0.48 switches per hour, while UnderPct changes only slightly from 0.081 to 0.085%. These findings indicate improved cost efficiency and controller stability across the evaluated workloads. The reported shortfall measures should be interpreted as capacity-level risk proxies rather than direct application-level service objectives, because they do not measure latency, throughput, request-error rates, or related end-to-end SLA outcomes.

Authors

Institutions

Publication Details

Journal
Scientific Reports
Published
2026-10-06
DOI
https://doi.org/10.1038/s41598-026-74671-5
Primary Topic
Cloud Computing and Resource Management
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Uncertainty-aware hybrid autoscaling with bias-corrected deep forecasting for cost-efficient capacity-shortfall control across cloud workload traces

A. Pirkhedri, Arman Kavoosi Ghafi, Mahdee Jodayree, Fatemeh Mokhtari et al.
Scientific Reports
Cloud Computing and Resource Management
article

Uncertainty-aware hybrid autoscaling with bias-corrected deep forecasting for cost-efficient capacity-shortfall control across cloud workload traces

A. Pirkhedri, Arman Kavoosi Ghafi, Mahdee Jodayree, Fatemeh Mokhtari, Puya Shaykholeslami
article en

Abstract

We present HybridTuned, an uncertainty-aware hybrid autoscaling framework designed to mitigate bias-induced capacity shortfalls when continuous workload forecasts are translated into discrete provisioning decisions. The primary evaluation uses the Google 2019 Cluster sample, aggregated at 5-minute intervals and partitioned chronologically into training, validation, and held-out test blocks. Point-forecasting performance is compared across LSTM, vanilla Transformer, and PatchTST models over five training seeds using Holm-adjusted Diebold–Mariano tests, while a quantile LSTM baseline is assessed separately using pinball loss and empirical coverage. HybridTuned combines median-based forecast-bias correction, a scale-normalized uncertainty margin derived from positive bias-corrected validation residuals, and a utilization-triggered reactive override. Controller parameters are selected exclusively through validation replay under an explicit capacity-shortfall feasibility constraint, and uncertainty in controller outcomes is quantified using paired moving-block bootstrap analysis. On the 1353-interval Google test trajectory, HybridTuned reduces replica-equivalent machine-hours from 5200.83 to 4,638.58, corresponding to a 10.81% reduction relative to the Reactive baseline. Both policies produce one shortage interval, yielding the same under-provisioning percentage (UnderPct) of 0.07391%. HybridTuned also reduces mean over-provisioning from 19.806 to 14.819 and switching frequency from 3.1308 to 0.2927 switches per hour. Cross-dataset evaluation on the GWA/Bitbrains traces yields an 8.0% cost reduction, decreases mean waste from 17.42 to 13.96, and reduces switching frequency from 2.41 to 0.48 switches per hour, while UnderPct changes only slightly from 0.081 to 0.085%. These findings indicate improved cost efficiency and controller stability across the evaluated workloads. The reported shortfall measures should be interpreted as capacity-level risk proxies rather than direct application-level service objectives, because they do not measure latency, throughput, request-error rates, or related end-to-end SLA outcomes.

Scientific Reports
Islamic Azad University, Tehran (IR), Islamic Azad University, Science and Research Branch (IR), University of Tehran (IR), Islamic Azad University Boroujerd Branch (IR), McMaster University (CA)
Openalex Percentile: Top 5%
Cloud Computing and Resource Management
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.