Nikhil Parkar, M. Marimuthu, Aditi Agale · Artificial Intelligence Review 2026 · 2026
DOI: 10.1007/s10462-026-11706-3
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Reinforcement learning (RL) has emerged as a promising approach to upgrade recommender systems, particularly in dynamic and evolving user environments. Conventional approaches such as collaborative filtering and matrix factorization are hampered by static modeling, data sparsity, and short-term thinking. RL, however, provides an approach that allows systems to learn and evolve in response to user interactions and optimize for immediate and long-term user engagement. This review examines the use of RL in recommender systems through the examination of over 60 works of literature spanning three main approaches: model-free, model-based, and hybrid approaches. This work also highlights several gaps in previous reviews, namely their lack of coverage of deep RL techniques and lack of systematic taxonomies for each of the three approaches. An important discovery in this area is that traditional metrics for offline evaluation such as Precision@K and NDCG cannot account for the long-term performance of policies, showing an inherent disconnect between offline evaluation and actual implementation. Multi-agent reinforcement learning (MARL) and offline RL are two of the most significant emergent areas of study, with the former allowing for complex multi-stage recommendation processes and the latter ensuring data efficiency without needing live interaction.
No comments yet — start the discussion below.