Xinhao Zhang · Applied and Computational Engineering 2026 · 2026
DOI: 10.54254/2755-2721/2026.37323
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
A fundamental problem in reinforcement learning is the trade-off between exploration and exploitation. The quality of the final policy is not the only factor that varies with different exploration strategies, so do learning speed and result stability. Based on Q-learning, this study compares three exploration strategies in a custom 5×5 GridWorld environment: fixed ε=0.10, fixed ε=0.30, and linearly decaying ε (1.00→0.05). A baseline experiment is first performed with slip_prob=0.10. The environmental noise is then extended to three levels, 0.00, 0.10, and 0.20, and the number of episodes required for the rolling 50-episode success rate to first reach 50% is used to measure learning speed. The results show that the fixed lower exploration rate learns faster overall under all three noise levels and demonstrates better stability across random seeds. Decaying exploration achieves a higher final success rate under low-noise conditions, but converges much more slowly in the early stage. The fixed higher exploration rate performs relatively poorly overall and has particular difficulty forming a stable policy in high-noise environments. These experiments show that exploration strategies should not be evaluated only by final success rate; environmental randomness, learning speed, and stability across random seeds should also be considered.
No comments yet — start the discussion below.