Microgrid energy management optimization using PPO with curriculum learning trained in virtual simulation environment

Abstract Microgrid energy management systems (EMS) require real-time, economically optimal, and safe dispatch strategies under high renewable uncertainty. While Deep Reinforcement Learning (DRL) offers promising online decision-making capabilities, standard DRL algorithms struggle with the “cold start” problem, slow convergence, and severe Sim-to-Real performance degradation when transferred from simulation to real-world operating conditions. To address these challenges, this paper proposes a novel Curriculum Learning-enhanced Proximal Policy Optimization (CL-PPO) framework for grid-connected microgrid energy management. By structuring a high-fidelity virtual environment into a three-stage progressive curriculum (Basic, Volatile, and Extreme), the agent systematically evolves from learning fundamental arbitrage rules to mastering robust dispatch under severe boundary conditions. Experimental results demonstrate that curriculum-guided knowledge transfer reduces the required convergence steps by approximately 38% compared with Baseline PPO. In typical-scenario evaluations, CL-PPO reduces daily operating costs by 7.1% relative to Baseline PPO and achieves a near-optimal 2.4% economic gap relative to the non-causal Mixed-Integer Linear Programming (MILP) oracle, requiring only 12.5 milliseconds per inference. Furthermore, an offline Sim-to-Real replay evaluation was conducted using approximately 90,000 previously unseen operational records from a real-world campus microgrid. In a replay environment parameterized according to the target microgrid’s energy-storage-system characteristics, CL-PPO demonstrated strong cross-domain robustness against measurement noise and operational disturbances, achieving an average total constraint violation rate of 1.2%. No policy-generated command was applied to the physical microgrid during this evaluation. Ultimately, this research enhances the interpretability of DRL policy evolution and provides a safety-aware and computationally efficient training framework with promising potential for future physical deployment, subject to further hardware-in-the-loop and on-site closed-loop validation.

Authors

Publication Details

Journal
Journal of Engineering and Applied Science
Published
2026-09-26
DOI
https://doi.org/10.1186/s44147-026-01236-8
Primary Topic
Microgrid Control and Optimization
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Microgrid energy management optimization using PPO with curriculum learning trained in virtual simulation environment

Yantao Li, Jiaying Li, Can Pei
Journal of Engineering and Applied Science
Microgrid Control and Optimization
article

Microgrid energy management optimization using PPO with curriculum learning trained in virtual simulation environment

Yantao Li, Jiaying Li, Can Pei
article en

Abstract

Abstract Microgrid energy management systems (EMS) require real-time, economically optimal, and safe dispatch strategies under high renewable uncertainty. While Deep Reinforcement Learning (DRL) offers promising online decision-making capabilities, standard DRL algorithms struggle with the “cold start” problem, slow convergence, and severe Sim-to-Real performance degradation when transferred from simulation to real-world operating conditions. To address these challenges, this paper proposes a novel Curriculum Learning-enhanced Proximal Policy Optimization (CL-PPO) framework for grid-connected microgrid energy management. By structuring a high-fidelity virtual environment into a three-stage progressive curriculum (Basic, Volatile, and Extreme), the agent systematically evolves from learning fundamental arbitrage rules to mastering robust dispatch under severe boundary conditions. Experimental results demonstrate that curriculum-guided knowledge transfer reduces the required convergence steps by approximately 38% compared with Baseline PPO. In typical-scenario evaluations, CL-PPO reduces daily operating costs by 7.1% relative to Baseline PPO and achieves a near-optimal 2.4% economic gap relative to the non-causal Mixed-Integer Linear Programming (MILP) oracle, requiring only 12.5 milliseconds per inference. Furthermore, an offline Sim-to-Real replay evaluation was conducted using approximately 90,000 previously unseen operational records from a real-world campus microgrid. In a replay environment parameterized according to the target microgrid’s energy-storage-system characteristics, CL-PPO demonstrated strong cross-domain robustness against measurement noise and operational disturbances, achieving an average total constraint violation rate of 1.2%. No policy-generated command was applied to the physical microgrid during this evaluation. Ultimately, this research enhances the interpretability of DRL policy evolution and provides a safety-aware and computationally efficient training framework with promising potential for future physical deployment, subject to further hardware-in-the-loop and on-site closed-loop validation.

Journal of Engineering and Applied ScienceVol. 73(1)
Affordable and clean energy
Openalex Percentile: Top 16%
Microgrid Control and Optimization
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Microgrid energy management optimization using PPO with curriculum learning trained in virtual simulation environment — Yantao Li, Jiaying Li, et al. · Journal of Engineering and Applied Science (2026) | TGRS Research Map | TGRS