Danil Gusak, Anna Volodkevich, Evgeny Frolov · ACM Conference on Recommender Systems (RecSys) 2026 · 2026
DOI: 10.1145/3773078.3841281
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
More test users should make offline evaluation more precise. For stochastic recommenders, however, paired tests over users can become overconfident. They compare two models, not whether one training procedure performs better on average than the other. Drawing conclusions about those procedures requires repeated runs, not more users. A variance model separating run-to-run and user-to-user variation gives the inflation factor \(1+U\sigma _a^2/\sigma _{\mathrm{id}}^2\): more users reduce uncertainty about the two models but not across seeds. With 15 SASRec runs on fixed sequential benchmarks, this inflation predicts how often user-level tests reject when comparing runs of the same procedure: rejection rates reach 80% on the 138K-user MovieLens-20M, while tests across runs remain near 5%. For near-ties in the commonly reported 1–2% range, the winner changes with the seed. We propose a lightweight seed-aware protocol using paired differences across runs and a variance-inflation diagnostic.
Last synced
This work has 2 recorded citations, but citing papers have not been linked locally yet.
No comments yet — start the discussion below.