Uncertainty-aware hybrid autoscaling with bias-corrected deep forecasting for cost-efficient capacity-shortfall control across cloud workload traces
We present HybridTuned, an uncertainty-aware hybrid autoscaling framework designed to mitigate bias-induced capacity shortfalls when continuous workload forecasts are translated into discrete provisioning decisions. The primary evaluation uses the Google 2019 Cluster sample, aggregated at 5-minute intervals and partitioned chronologically into training, validation, and held-out test blocks. Point-forecasting performance is compared across LSTM, vanilla Transformer, and PatchTST models over five training seeds using Holm-adjusted Diebold–Mariano tests, while a quantile LSTM baseline is assessed separately using pinball loss and empirical coverage. HybridTuned combines median-based forecast-bias correction, a scale-normalized uncertainty margin derived from positive bias-corrected validation residuals, and a utilization-triggered reactive override. Controller parameters are selected exclusively through validation replay under an explicit capacity-shortfall feasibility constraint, and uncertainty in controller outcomes is quantified using paired moving-block bootstrap analysis. On the 1353-interval Google test trajectory, HybridTuned reduces replica-equivalent machine-hours from 5200.83 to 4,638.58, corresponding to a 10.81% reduction relative to the Reactive baseline. Both policies produce one shortage interval, yielding the same under-provisioning percentage (UnderPct) of 0.07391%. HybridTuned also reduces mean over-provisioning from 19.806 to 14.819 and switching frequency from 3.1308 to 0.2927 switches per hour. Cross-dataset evaluation on the GWA/Bitbrains traces yields an 8.0% cost reduction, decreases mean waste from 17.42 to 13.96, and reduces switching frequency from 2.41 to 0.48 switches per hour, while UnderPct changes only slightly from 0.081 to 0.085%. These findings indicate improved cost efficiency and controller stability across the evaluated workloads. The reported shortfall measures should be interpreted as capacity-level risk proxies rather than direct application-level service objectives, because they do not measure latency, throughput, request-error rates, or related end-to-end SLA outcomes.
Authors
- A. Pirkhedri (ORCID: https://orcid.org/0000-0003-3752-3852)
- Arman Kavoosi Ghafi (ORCID: https://orcid.org/0009-0001-8907-449X)
- Mahdee Jodayree
- Fatemeh Mokhtari (ORCID: https://orcid.org/0009-0004-2981-0450)
- Puya Shaykholeslami
Institutions
- Islamic Azad University, Tehran (IR)
- Islamic Azad University, Science and Research Branch (IR)
- University of Tehran (IR)
- Islamic Azad University Boroujerd Branch (IR)
- McMaster University (CA)
Publication Details
- Journal
- Scientific Reports
- Published
- 2026-10-06
- DOI
- https://doi.org/10.1038/s41598-026-74671-5
- Primary Topic
- Cloud Computing and Resource Management
- Type
- article
- Field-Weighted Citation Impact
- 0.00