Anusuya S, Anandapadmanabhan K. · Mapana Journal of Sciences 2026 · 2026
DOI: 10.12723/mjs.78.8
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
World Wide Web is an immense repository of information, the sheer volume of data makes it challenging to locate specific information easily. Recommendation Systems (RS) are designed to alleviate this issue by identifying similar items and customers based on their behaviour, subsequently suggesting items tailored to individual preferences. The RS involves RL (Reinforcement Learning)-based RS approaches used to overcome existing limitations by fusing RL with self-supervised sequential learning. However, these approaches often suffer from biases in estimating Q-values. It exclusively depends To rectify this, a novel strategy incorporating Supervised Negative Q-learning and Supervised Advantage Actor-Critic has been introduced. This work proposes a solution to enhance stability by merging RL with sequential modelling, contrastive-based objectives, and negative sampling methods, alongwith contrastive learning and conservative Q-learning. Together, these components improve performance and stability. The proposed method validates its effectiveness through empirical results obtained from real-world datasets.
No comments yet — start the discussion below.