Research on storage early warning and scheduling for cloud data centers based on spatial-temporal graph atlas and dynamic extreme value theory
Abstract Under cloud-native microservice architectures, storage workloads exhibit complex spatial-temporal cascading effects and significant non-stationary heavy-tail characteristics. Traditional passive operations and maintenance rules are inadequate to cope with concept drift induced by high concurrency, and are prone to large-scale alert storms and Service Level Agreement (SLA) violations. To address these pain points, this paper proposes a joint prediction-alerting and scheduling architecture named STG-SPOT(Spatial-Temporal Graph with Streaming Peak-Over Threshold (SPOT)), aiming to break down operational silos in data centers and build an end-to-end intelligent autonomous closed loop. First, at the perception layer, an improved Graph Attention Network (GAT) and Dilated Causal Convolution (TCN) are fused to precisely extract the spatial spillover effects of microservice topology and long-term temporal evolution trends. Second, at the alerting layer, Extreme Value Theory (EVT) and continuous sliding window maximum likelihood estimation are introduced to dynamically track the tail distribution of workloads in an unsupervised manner, deriving adaptive dynamic alert boundaries. Finally, at the execution layer, an operations research model targeting Total Cost of Ownership (TCO) minimization is constructed, and an improved Genetic Algorithm (GA) is employed to output Pareto-optimal storage tiering strategies. Experimental results on large-scale industrial cluster datasets from Alibaba and Google demonstrate that STG-SPOT reduces prediction error by approximately $$9.5\\%$$ compared to the best baseline, achieves a maximum alerting F1-Score of $$0.947$$ , and realizes cost savings exceeding $$46\\%$$ while keeping the SLA violation rate below $$1\\%$$ . This architecture provides a highly efficient engineering paradigm for modern computing centers to handle heavy-tail bursts and achieve high-availability scheduling.
Authors
- Chen Wang
- Zepeng Wen
- Shuai Xu
Institutions
- China Academy of Engineering Physics (CN)
Publication Details
- Journal
- Journal of Cloud Computing Advances Systems and Applications
- Published
- 2026-09-14
- DOI
- https://doi.org/10.1186/s13677-026-00983-6
- Primary Topic
- Software System Performance and Reliability
- Type
- article
- Field-Weighted Citation Impact
- 0.00