A Swarm Optimization-Based Deep Reinforcement Learning Approach for Nonlinear Control Systems with Shapley Value Interpretability

This study introduces the Proposed Optimization Algorithm (POA), a swarmbasedhybrid combining the Gorilla Troops Optimizer (GTO) and the Artificial Bee Colony(ABC) algorithm, to enhance reward maximization in control tasks. Neural networks weretrained for both simple (pendulum) and complex (bipedal walker) environments. The POAalgorithm was run 10 times, and based on the resulting median values, the best rewardscores achieved were −117.771 in the pendulum environment and 30.936 in the bipedalwalker environment. These reward values indicate a 0.244 improvement for the pendulumenvironment and a 39.569 improvement for the bipedal walker compared to the closestcompetitors (GTO). While there was no statistically significant difference between GTO andPOA in the pendulum task, POA performed significantly better than all other algorithmsin the bipedal walker environment (p < 0.05). To address the “black box” nature ofreinforcement learning, the study integrated Shapley Value Theory for post-training analysis.This explainable AI (XAI) approach identified angular velocity as the primary driver oftorque in the pendulum task and quantified the importance of observation parameters forthe bipedal walker. The results provide both a high-performing optimization framework anda robust method for interpreting neural network decision-making in robotic control systems.

Authors

Institutions

Publication Details

Journal
Journal of Advanced Research in Natural and Applied Sciences
Published
2026-09-30
DOI
https://doi.org/10.28979/jarnas.1949124
Primary Topic
Reinforcement Learning in Robotics
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

A Swarm Optimization-Based Deep Reinforcement Learning Approach for Nonlinear Control Systems with Shapley Value Interpretability

Mustafa Can Bingöl
Journal of Advanced Research in Natural and Applied Sciences
Reinforcement Learning in Robotics
article

A Swarm Optimization-Based Deep Reinforcement Learning Approach for Nonlinear Control Systems with Shapley Value Interpretability

Mustafa Can Bingöl
article en

Abstract

This study introduces the Proposed Optimization Algorithm (POA), a swarmbasedhybrid combining the Gorilla Troops Optimizer (GTO) and the Artificial Bee Colony(ABC) algorithm, to enhance reward maximization in control tasks. Neural networks weretrained for both simple (pendulum) and complex (bipedal walker) environments. The POAalgorithm was run 10 times, and based on the resulting median values, the best rewardscores achieved were −117.771 in the pendulum environment and 30.936 in the bipedalwalker environment. These reward values indicate a 0.244 improvement for the pendulumenvironment and a 39.569 improvement for the bipedal walker compared to the closestcompetitors (GTO). While there was no statistically significant difference between GTO andPOA in the pendulum task, POA performed significantly better than all other algorithmsin the bipedal walker environment (p < 0.05). To address the “black box” nature ofreinforcement learning, the study integrated Shapley Value Theory for post-training analysis.This explainable AI (XAI) approach identified angular velocity as the primary driver oftorque in the pendulum task and quantified the importance of observation parameters forthe bipedal walker. The results provide both a high-performing optimization framework anda robust method for interpreting neural network decision-making in robotic control systems.

Journal of Advanced Research in Natural and Applied SciencesVol. 12(3)
Burdur Mehmet Akif Ersoy Üniversitesi (TR)
Peace, Justice and strong institutions
Openalex Percentile: Top 9%
Reinforcement Learning in Robotics
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.