Yu Xie, Xing Kai Ren, Ying Qi, Ping Yang, Yao Hu · · 2026
DOI: 10.1145/3773078.3831745
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Generative recommenders increasingly train Semantic-ID models to autoregressively generate length-L slates, aligned to user feedback by reinforcement learning. Gradient-Bounded Policy Optimization (GBPO) stabilizes this stage with a BCE-stable asymmetric gradient bound, but its clipping boundary is computed per sample and is therefore blind to three recommendation-specific signals—rare-positive value, low-diversity negative slates, and multi-objective reward resolution—a limitation we call the Static Boundary Bottleneck. We propose SAGE (Sequence-level Adaptive Gradient Evolution), which keeps the OneRec-style architecture and GBPO’s SFT-then-RLHF interface but replaces the static boundary with one slate-aware adaptive clipped surrogate: a sequence-level geometric-mean ratio sets the update to the generated slate, positive/negative adaptive bounds condition clipping on advantage sign and slate diversity, and a decouple-then-aggregate advantage preserves multi-objective resolution. Changing only the optimizer, OneRec-SAGE preserves top-K accuracy while improving offline cold-start recall by +89.9% to +101% and diversity by approximately +11% on Amazon, with consistent beyond-accuracy gains on RecIF-Bench—evidence that making the clipping boundary slate-aware, rather than changing the retrieval backbone, is the primary lever for rare-item exposure and list diversity on these offline proxies.
No comments yet — start the discussion below.