Muhammad Fakhri Loebis, Andry Chowanda · Bulletin of Electrical Engineering and Informatics 2026 · 2026
DOI: 10.11591/eei.v15i5.11864
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
This study presents a comparative analysis of proximal policy optimization (PPO) and soft actor-critic (SAC) for training autonomous delivery agents in high-fidelity 3D environments using Unity ML-Agents. Both algorithms were evaluated with identical hyperparameters and reward functions across five independent runs to ensure statistical robustness. Results show that PPO achieves 31.4% higher final reward (6658.82 versus 5067.27) and superior policy improvement consistency (ratio 2.00 versus 0.31). However, PPO exhibits greater reward variability (coefficient of variation (CV) 0.533 versus 0.107) and slower convergence (33 versus 6 steps). SAC demonstrates faster initial convergence and superior stability with lower performance variance. These findings indicate that PPO is preferable for applications prioritizing maximum final performance in stable environments, while SAC is more suitable for tasks requiring adaptability and consistent performance under dynamic conditions. This study provides practical guidance for researchers implementing reinforcement learning (RL) in Unity-based simulations for autonomous systems.
No comments yet — start the discussion below.