Real-time optimization and comparative performance evaluation of ST-SMC, MFAC, and an online adaptive reinforcement learning controller using the Mayfly algorithm
In this study, the real-time performances of reinforcement learning (RL), model-free adaptive control (MFAC), and super-twisting sliding mode control (ST-SMC) methods were experimentally compared. The implemented RL controller employs a deterministic direct online adaptive policy-update structure rather than a conventional Q-learning, actor–critic, or value-function-based architecture, and it does not require stochastic exploration or offline training. Conventional fixed-parameter control schemes may exhibit performance degradation under nonlinear and operating-point-dependent dynamics, while advanced robust, adaptive, and learning-based approaches involve different trade-offs in terms of disturbance rejection, adaptability, and transient performance, motivating their systematic evaluation under identical real-time conditions. The Mayfly optimization algorithm was employed to determine the controller parameters. The algorithm was implemented in real time on an STM32F407-based experimental test system incorporating a boost DC–DC converter, and the optimal parameters for each control method were determined under the same experimental setup and evaluation conditions. Using the optimized parameters, all controllers achieved successful reference tracking with low steady-state error. RL provided the most balanced overall performance, whereas MFAC stood out with its fast rise response and ST-SMC exhibited a stable but relatively slower nominal response. The extended quantitative robustness analysis showed that MFAC provided low accumulated tracking error and fast recovery in several reference and disturbance transitions, ST-SMC most effectively limited voltage deviations during load disturbances and input-voltage reductions, and RL consistently required the smallest control-effort variation during the load- and input-voltage-change tests while maintaining strong noise robustness. The overall results indicate that no single control method achieved absolute superiority under all test conditions. Considering the adopted objective function together with the nominal and measurement-noise results, the RL controller provided the most balanced overall performance, while the disturbance tests revealed complementary advantages of MFAC and ST-SMC. Under nominal operation, RL achieved an objective-function value of 3.6541, a settling time of 0.0960 s, and a maximum overshoot of 0.3091%, compared with 5.2478, 0.1915s, and 2.1049% for MFAC and 7.8905, 0.2711 s, and 1.2297% for ST-SMC, respectively. Furthermore, the Mayfly algorithm was shown to be an effective approach for the real-time experimental optimization and comparative evaluation of structurally different control methods.
Authors
- Dogan Can Samuk
- Ahmet Yuksel
Institutions
- Recep Tayyip Erdoğan University (TR)
- Türkiye Bilimsel ve Teknolojik Araştırma Kurumu (TR)
Publication Details
- Journal
- Scientific Reports
- Published
- 2026-09-21
- DOI
- https://doi.org/10.1038/s41598-026-72390-5
- Primary Topic
- Adaptive Dynamic Programming Control
- Type
- article
- Field-Weighted Citation Impact
- 0.00