Mustafa Can Bingöl · Journal of Advanced Research in Natural and Applied Sciences 2026 · 2026
DOI: 10.28979/jarnas.1949124
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
This study introduces the Proposed Optimization Algorithm (POA), a swarmbasedhybrid combining the Gorilla Troops Optimizer (GTO) and the Artificial Bee Colony(ABC) algorithm, to enhance reward maximization in control tasks. Neural networks weretrained for both simple (pendulum) and complex (bipedal walker) environments. The POAalgorithm was run 10 times, and based on the resulting median values, the best rewardscores achieved were −117.771 in the pendulum environment and 30.936 in the bipedalwalker environment. These reward values indicate a 0.244 improvement for the pendulumenvironment and a 39.569 improvement for the bipedal walker compared to the closestcompetitors (GTO). While there was no statistically significant difference between GTO andPOA in the pendulum task, POA performed significantly better than all other algorithmsin the bipedal walker environment (p < 0.05). To address the “black box” nature ofreinforcement learning, the study integrated Shapley Value Theory for post-training analysis.This explainable AI (XAI) approach identified angular velocity as the primary driver oftorque in the pendulum task and quantified the importance of observation parameters forthe bipedal walker. The results provide both a high-performing optimization framework anda robust method for interpreting neural network decision-making in robotic control systems.
No comments yet — start the discussion below.