A fast risk assessment method for power balance under extreme weather conditions based on deep reinforcement learning with a coupled attention mechanism

Even if persistent stagnant weather does not directly damage power grid equipment, it may lead to prolonged insufficient output of renewable energy sources and further trigger power balance risks. This paper proposes a rapid risk assessment framework that comprehensively incorporates meteorological sensitivity models, time-varying component failure probabilities, a Denoising Variational Autoencoder (DVAE) for generating coupled wind-photovoltaic-temperature scenarios, and a Deep Deterministic Policy Gradient (DDPG) model embedded with the feature token attention mechanism.The 72-hour assessment task is formulated as an offline Markov Decision Process (MDP): the state space covers meteorological conditions, power sources, power grid topology, load profiles, and risk features of the previous time step; continuous actions are defined as risk indicator vectors; the reward function is designed to penalize deviations between model outputs and PSD-BPA benchmarks, temporal inconsistency, and violations of physical constraints.To distinguish the sequential reinforcement learning formulation from direct regression, a feedforward Multi-Layer Perceptron (MLP) is introduced as the non-reinforcement learning baseline. Four metrics, namely Mean Absolute Error (MAE), Root Mean Square Error (RMSE), coefficient of determination (R 2 ), and the 95th percentile of absolute error (P95AE), are adopted to evaluate all models on an independent scenario-level test set.Verified on the planned power grid of Northeast China in 2025 as the test case, the proposed DDPG-Attention achieves MAE of 0.0310, RMSE of 0.0389, R 2 of 0.954 and P95AE of 0.0760 on the independent test set. The average time consumption for a single risk assessment is approximately 0.12 seconds, yielding an acceleration ratio of multiple orders of magnitude compared with the traditional time-sequential Monte Carlo simulation based on PSD-BPA. This study provides an efficient and accurate novel technical paradigm for the rapid evaluation of power supply-demand imbalance risks in power systems under extreme weather conditions.

Authors

Institutions

Publication Details

Journal
PLoS ONE
Published
2026-10-08
DOI
https://doi.org/10.1371/journal.pone.0359101
Primary Topic
Power System Reliability and Maintenance
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

A fast risk assessment method for power balance under extreme weather conditions based on deep reinforcement learning with a coupled attention mechanism

Baoju Li, Cong Yu, Hang Wang, Meiyue Xu et al.
PLoS ONE
Power System Reliability and Maintenance
article

A fast risk assessment method for power balance under extreme weather conditions based on deep reinforcement learning with a coupled attention mechanism

Baoju Li, Cong Yu, Hang Wang, Meiyue Xu, Niu Hexiang, Shouheng Sun, Li Yiming
article en

Abstract

Even if persistent stagnant weather does not directly damage power grid equipment, it may lead to prolonged insufficient output of renewable energy sources and further trigger power balance risks. This paper proposes a rapid risk assessment framework that comprehensively incorporates meteorological sensitivity models, time-varying component failure probabilities, a Denoising Variational Autoencoder (DVAE) for generating coupled wind-photovoltaic-temperature scenarios, and a Deep Deterministic Policy Gradient (DDPG) model embedded with the feature token attention mechanism.The 72-hour assessment task is formulated as an offline Markov Decision Process (MDP): the state space covers meteorological conditions, power sources, power grid topology, load profiles, and risk features of the previous time step; continuous actions are defined as risk indicator vectors; the reward function is designed to penalize deviations between model outputs and PSD-BPA benchmarks, temporal inconsistency, and violations of physical constraints.To distinguish the sequential reinforcement learning formulation from direct regression, a feedforward Multi-Layer Perceptron (MLP) is introduced as the non-reinforcement learning baseline. Four metrics, namely Mean Absolute Error (MAE), Root Mean Square Error (RMSE), coefficient of determination (R 2 ), and the 95th percentile of absolute error (P95AE), are adopted to evaluate all models on an independent scenario-level test set.Verified on the planned power grid of Northeast China in 2025 as the test case, the proposed DDPG-Attention achieves MAE of 0.0310, RMSE of 0.0389, R 2 of 0.954 and P95AE of 0.0760 on the independent test set. The average time consumption for a single risk assessment is approximately 0.12 seconds, yielding an acceleration ratio of multiple orders of magnitude compared with the traditional time-sequential Monte Carlo simulation based on PSD-BPA. This study provides an efficient and accurate novel technical paradigm for the rapid evaluation of power supply-demand imbalance risks in power systems under extreme weather conditions.

PLoS ONEVol. 21(10)
Northeast Electric Power University (CN), Jilin Electric Power Research Institute (China) (CN), State Grid Jilin Electric Power (China) (CN)
Openalex Percentile: Top 11%
Power System Reliability and Maintenance
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.