
Arwa S. Bazmalah, Noorfazila Binti Kamal, Kalaivani Chellappan, Asraf Mohamed Moubark, Arwa S. Bazmalah · Scientific Reports 2026 · 2026
DOI: 10.1038/s41598-026-73214-2
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Q-learning (QL) faces a challenge in balancing exploration and exploitation. QL typically employs an ε-greedy policy. However, the ε-greedy policy leads to slow convergence and excessive resource consumption on Field-Programmable Gate Arrays (FPGAs) due to unnecessary exploration. This paper proposes a hardware Q-learning architecture that integrates the Boltzmann policy using fixed-point representation. The design is implemented on the Genesys 2 Kintex7 (XC7K325T-2FFG900C) FPGA. Utilizing the Boltzmann policy's ability to prioritize optimal actions, the present design significantly reduces the number of iterations required to update states. This reduction results in faster convergence and a more efficient resource allocation strategy. Experimental results demonstrate improved FPGA resource efficiency. For 16-bit implementations, the proposed design limits the usage of Look-Up Tables (LUTs), LUTRAMs, Flip-Flops (FFs), and Block Random Access Memories (BRAMs) to just 0.86%, 0.06%, 0.27%, and 0.79% of the total available hardware resources, respectively. Similar for 32-bit implementations, the usage of LUTs, LUTRAMs, FFs, and BRAMs is restricted to 1.26%, 0.01%, 0.38%, and 2.02%, respectively. Moreover, the proposed design achieves power efficiency improvements of 46% for 16-bit configurations and 28% for 32-bit configurations. Based on these results, the proposed Boltzmann-based QL architecture effectively optimizes FPGA resource utilization while enhancing overall system performance.
No comments yet — start the discussion below.