Research on storage early warning and scheduling for cloud data centers based on spatial-temporal graph atlas and dynamic extreme value theory

Abstract Under cloud-native microservice architectures, storage workloads exhibit complex spatial-temporal cascading effects and significant non-stationary heavy-tail characteristics. Traditional passive operations and maintenance rules are inadequate to cope with concept drift induced by high concurrency, and are prone to large-scale alert storms and Service Level Agreement (SLA) violations. To address these pain points, this paper proposes a joint prediction-alerting and scheduling architecture named STG-SPOT(Spatial-Temporal Graph with Streaming Peak-Over Threshold (SPOT)), aiming to break down operational silos in data centers and build an end-to-end intelligent autonomous closed loop. First, at the perception layer, an improved Graph Attention Network (GAT) and Dilated Causal Convolution (TCN) are fused to precisely extract the spatial spillover effects of microservice topology and long-term temporal evolution trends. Second, at the alerting layer, Extreme Value Theory (EVT) and continuous sliding window maximum likelihood estimation are introduced to dynamically track the tail distribution of workloads in an unsupervised manner, deriving adaptive dynamic alert boundaries. Finally, at the execution layer, an operations research model targeting Total Cost of Ownership (TCO) minimization is constructed, and an improved Genetic Algorithm (GA) is employed to output Pareto-optimal storage tiering strategies. Experimental results on large-scale industrial cluster datasets from Alibaba and Google demonstrate that STG-SPOT reduces prediction error by approximately $$9.5\\%$$ compared to the best baseline, achieves a maximum alerting F1-Score of $$0.947$$ , and realizes cost savings exceeding $$46\\%$$ while keeping the SLA violation rate below $$1\\%$$ . This architecture provides a highly efficient engineering paradigm for modern computing centers to handle heavy-tail bursts and achieve high-availability scheduling.

Authors

Institutions

Publication Details

Journal
Journal of Cloud Computing Advances Systems and Applications
Published
2026-09-14
DOI
https://doi.org/10.1186/s13677-026-00983-6
Primary Topic
Software System Performance and Reliability
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Research on storage early warning and scheduling for cloud data centers based on spatial-temporal graph atlas and dynamic extreme value theory

Chen Wang, Zepeng Wen, Shuai Xu
Journal of Cloud Computing Advances Systems and Applications
Software System Performance and Reliability
article

Research on storage early warning and scheduling for cloud data centers based on spatial-temporal graph atlas and dynamic extreme value theory

Chen Wang, Zepeng Wen, Shuai Xu
article en

Abstract

Abstract Under cloud-native microservice architectures, storage workloads exhibit complex spatial-temporal cascading effects and significant non-stationary heavy-tail characteristics. Traditional passive operations and maintenance rules are inadequate to cope with concept drift induced by high concurrency, and are prone to large-scale alert storms and Service Level Agreement (SLA) violations. To address these pain points, this paper proposes a joint prediction-alerting and scheduling architecture named STG-SPOT(Spatial-Temporal Graph with Streaming Peak-Over Threshold (SPOT)), aiming to break down operational silos in data centers and build an end-to-end intelligent autonomous closed loop. First, at the perception layer, an improved Graph Attention Network (GAT) and Dilated Causal Convolution (TCN) are fused to precisely extract the spatial spillover effects of microservice topology and long-term temporal evolution trends. Second, at the alerting layer, Extreme Value Theory (EVT) and continuous sliding window maximum likelihood estimation are introduced to dynamically track the tail distribution of workloads in an unsupervised manner, deriving adaptive dynamic alert boundaries. Finally, at the execution layer, an operations research model targeting Total Cost of Ownership (TCO) minimization is constructed, and an improved Genetic Algorithm (GA) is employed to output Pareto-optimal storage tiering strategies. Experimental results on large-scale industrial cluster datasets from Alibaba and Google demonstrate that STG-SPOT reduces prediction error by approximately $$9.5\%$$ compared to the best baseline, achieves a maximum alerting F1-Score of $$0.947$$ , and realizes cost savings exceeding $$46\%$$ while keeping the SLA violation rate below $$1\%$$ . This architecture provides a highly efficient engineering paradigm for modern computing centers to handle heavy-tail bursts and achieve high-availability scheduling.

Journal of Cloud Computing Advances Systems and Applications
China Academy of Engineering Physics (CN)
Openalex Percentile: Top 8%
Software System Performance and Reliability
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.