Real-time optimization and comparative performance evaluation of ST-SMC, MFAC, and an online adaptive reinforcement learning controller using the Mayfly algorithm

In this study, the real-time performances of reinforcement learning (RL), model-free adaptive control (MFAC), and super-twisting sliding mode control (ST-SMC) methods were experimentally compared. The implemented RL controller employs a deterministic direct online adaptive policy-update structure rather than a conventional Q-learning, actor–critic, or value-function-based architecture, and it does not require stochastic exploration or offline training. Conventional fixed-parameter control schemes may exhibit performance degradation under nonlinear and operating-point-dependent dynamics, while advanced robust, adaptive, and learning-based approaches involve different trade-offs in terms of disturbance rejection, adaptability, and transient performance, motivating their systematic evaluation under identical real-time conditions. The Mayfly optimization algorithm was employed to determine the controller parameters. The algorithm was implemented in real time on an STM32F407-based experimental test system incorporating a boost DC–DC converter, and the optimal parameters for each control method were determined under the same experimental setup and evaluation conditions. Using the optimized parameters, all controllers achieved successful reference tracking with low steady-state error. RL provided the most balanced overall performance, whereas MFAC stood out with its fast rise response and ST-SMC exhibited a stable but relatively slower nominal response. The extended quantitative robustness analysis showed that MFAC provided low accumulated tracking error and fast recovery in several reference and disturbance transitions, ST-SMC most effectively limited voltage deviations during load disturbances and input-voltage reductions, and RL consistently required the smallest control-effort variation during the load- and input-voltage-change tests while maintaining strong noise robustness. The overall results indicate that no single control method achieved absolute superiority under all test conditions. Considering the adopted objective function together with the nominal and measurement-noise results, the RL controller provided the most balanced overall performance, while the disturbance tests revealed complementary advantages of MFAC and ST-SMC. Under nominal operation, RL achieved an objective-function value of 3.6541, a settling time of 0.0960 s, and a maximum overshoot of 0.3091%, compared with 5.2478, 0.1915s, and 2.1049% for MFAC and 7.8905, 0.2711 s, and 1.2297% for ST-SMC, respectively. Furthermore, the Mayfly algorithm was shown to be an effective approach for the real-time experimental optimization and comparative evaluation of structurally different control methods.

Authors

Institutions

Publication Details

Journal
Scientific Reports
Published
2026-09-21
DOI
https://doi.org/10.1038/s41598-026-72390-5
Primary Topic
Adaptive Dynamic Programming Control
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Real-time optimization and comparative performance evaluation of ST-SMC, MFAC, and an online adaptive reinforcement learning controller using the Mayfly algorithm

Dogan Can Samuk, Ahmet Yuksel
Scientific Reports
Adaptive Dynamic Programming Control
article

Real-time optimization and comparative performance evaluation of ST-SMC, MFAC, and an online adaptive reinforcement learning controller using the Mayfly algorithm

Dogan Can Samuk, Ahmet Yuksel
article en

Abstract

In this study, the real-time performances of reinforcement learning (RL), model-free adaptive control (MFAC), and super-twisting sliding mode control (ST-SMC) methods were experimentally compared. The implemented RL controller employs a deterministic direct online adaptive policy-update structure rather than a conventional Q-learning, actor–critic, or value-function-based architecture, and it does not require stochastic exploration or offline training. Conventional fixed-parameter control schemes may exhibit performance degradation under nonlinear and operating-point-dependent dynamics, while advanced robust, adaptive, and learning-based approaches involve different trade-offs in terms of disturbance rejection, adaptability, and transient performance, motivating their systematic evaluation under identical real-time conditions. The Mayfly optimization algorithm was employed to determine the controller parameters. The algorithm was implemented in real time on an STM32F407-based experimental test system incorporating a boost DC–DC converter, and the optimal parameters for each control method were determined under the same experimental setup and evaluation conditions. Using the optimized parameters, all controllers achieved successful reference tracking with low steady-state error. RL provided the most balanced overall performance, whereas MFAC stood out with its fast rise response and ST-SMC exhibited a stable but relatively slower nominal response. The extended quantitative robustness analysis showed that MFAC provided low accumulated tracking error and fast recovery in several reference and disturbance transitions, ST-SMC most effectively limited voltage deviations during load disturbances and input-voltage reductions, and RL consistently required the smallest control-effort variation during the load- and input-voltage-change tests while maintaining strong noise robustness. The overall results indicate that no single control method achieved absolute superiority under all test conditions. Considering the adopted objective function together with the nominal and measurement-noise results, the RL controller provided the most balanced overall performance, while the disturbance tests revealed complementary advantages of MFAC and ST-SMC. Under nominal operation, RL achieved an objective-function value of 3.6541, a settling time of 0.0960 s, and a maximum overshoot of 0.3091%, compared with 5.2478, 0.1915s, and 2.1049% for MFAC and 7.8905, 0.2711 s, and 1.2297% for ST-SMC, respectively. Furthermore, the Mayfly algorithm was shown to be an effective approach for the real-time experimental optimization and comparative evaluation of structurally different control methods.

Scientific Reports
Recep Tayyip Erdoğan University (TR), Türkiye Bilimsel ve Teknolojik Araştırma Kurumu (TR)
Openalex Percentile: Top 9%
Adaptive Dynamic Programming Control
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.