Microgrid energy management optimization using PPO with curriculum learning trained in virtual simulation environment
Abstract Microgrid energy management systems (EMS) require real-time, economically optimal, and safe dispatch strategies under high renewable uncertainty. While Deep Reinforcement Learning (DRL) offers promising online decision-making capabilities, standard DRL algorithms struggle with the “cold start” problem, slow convergence, and severe Sim-to-Real performance degradation when transferred from simulation to real-world operating conditions. To address these challenges, this paper proposes a novel Curriculum Learning-enhanced Proximal Policy Optimization (CL-PPO) framework for grid-connected microgrid energy management. By structuring a high-fidelity virtual environment into a three-stage progressive curriculum (Basic, Volatile, and Extreme), the agent systematically evolves from learning fundamental arbitrage rules to mastering robust dispatch under severe boundary conditions. Experimental results demonstrate that curriculum-guided knowledge transfer reduces the required convergence steps by approximately 38% compared with Baseline PPO. In typical-scenario evaluations, CL-PPO reduces daily operating costs by 7.1% relative to Baseline PPO and achieves a near-optimal 2.4% economic gap relative to the non-causal Mixed-Integer Linear Programming (MILP) oracle, requiring only 12.5 milliseconds per inference. Furthermore, an offline Sim-to-Real replay evaluation was conducted using approximately 90,000 previously unseen operational records from a real-world campus microgrid. In a replay environment parameterized according to the target microgrid’s energy-storage-system characteristics, CL-PPO demonstrated strong cross-domain robustness against measurement noise and operational disturbances, achieving an average total constraint violation rate of 1.2%. No policy-generated command was applied to the physical microgrid during this evaluation. Ultimately, this research enhances the interpretability of DRL policy evolution and provides a safety-aware and computationally efficient training framework with promising potential for future physical deployment, subject to further hardware-in-the-loop and on-site closed-loop validation.
Authors
- Yantao Li
- Jiaying Li
- Can Pei
Publication Details
- Journal
- Journal of Engineering and Applied Science
- Published
- 2026-09-26
- DOI
- https://doi.org/10.1186/s44147-026-01236-8
- Primary Topic
- Microgrid Control and Optimization
- Type
- article
- Field-Weighted Citation Impact
- 0.00