Barrier-Function-Constrained Residual RL-Based Current Control for LCL-Filtered NPC Inverters under Weak-Grid Operation
As renewable energy penetration continues to grow, grid-connected inverters are increasingly required to operate reliably under weak-grid conditions, where low short-circuit ratios, changing grid impedances, and rapid disturbances reveal the limitations of conventional fixed-gain PI current controllers. Reinforcement learning (RL) offers a more adaptive control approach; however, its application to fast inner current loops remains limited by stringent safety requirements and practical rediness challenges. This paper presents a control framework that combines a model-free Proximal Policy Optimization (PPO)-based residual reinforcement learning controller with a model-based Control Lyapunov Function–Control Barrier Function (CLF–CBF) quadratic-program (QP) safety filter for an LCL-filtered three-phase neutral-point-clamped (NPC) grid-connected inverter. Although residual RL and CLF–CBF safety methods have been explored independently, this work integrates a model-free PPO-based residual controller with a model-based CLF–CBF safety filter within a unified architecture for fast inner-loop current control under varying grid-strength and operating conditions. Instead of replacing the conventional PI controller, the residual policy generates bounded corrective voltage commands that complement a fixed PI baseline, while the CLF–CBF–QP safety filter provides analytical enforcement of stability and operational constraints independently of the learned policy. A two-stage training strategy, consisting of unconstrained warm-start followed by safety-constrained fine-tuning, enables efficient policy learning while maintaining safe operation. Evaluated across eight representative operating scenarios with five independent trials per condition, the proposed framework achieves a maximum total harmonic distortion (THD) reduction of 49.7% under extreme thermal stress and 44.2% under nominal operation, with an average THD reduction of 25.7% across all scenarios. The framework also reduces aggregate constraint violations by 35.7% across the evaluated conditions, indicating that the integration of model-free residual reinforcement learning with a model-based CLF–CBF safety layer can improve harmonic performance while providing model-based supervision of current and voltage operating constraints.
Authors
- Shameem Ahmad (ORCID: https://orcid.org/0000-0002-6602-2064)
- Chowdhury Akram Hossain (ORCID: https://orcid.org/0000-0002-8769-2833)
- Md. Rifat Hazari (ORCID: https://orcid.org/0000-0002-1398-3013)
- Jabala Nur Fahima
- Mohammad Abdul Mannan (ORCID: https://orcid.org/0000-0002-4583-4610)
- Kayes Hasan
- Emanuele Ogliari (ORCID: https://orcid.org/0000-0002-2106-0374)
Institutions
- American International University-Bangladesh (BD)
- BRAC University (BD)
- RMIT University (AU)
- Politecnico di Milano (IT)
Publication Details
- Journal
- Electric Power Systems Research
- Published
- 2026-09-18
- DOI
- https://doi.org/10.1016/j.epsr.2026.114214
- Primary Topic
- Microgrid Control and Optimization
- Type
- article
- Field-Weighted Citation Impact
- 0.00