Danil Gusak, Evgeny Frolov · ACM Conference on Recommender Systems (RecSys) 2026 · 2026
DOI: 10.1145/3773078.3841280
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Temporal recommender benchmarks are often specified as “5-core plus temporal split,” but this phrase does not define a unique evaluation protocol. We show that iterative p-core filtering and temporal splitting do not commute: when the full log is filtered before the split, users and items can survive because of validation or test interactions unavailable during model fitting. The resulting fitted train split may therefore violate the same p-core property claimed for the benchmark. Across six timestamped datasets, fitted-train violations at p = 5 reach 38.2% of users and 22.3% of items. These protocol choices also reshape the evaluated population, targets, and candidate universe, and can change conclusions: across eight recommenders, the top-ranked model changes in 8 of 14 matched p-core settings with p > 0, and in 5 of 10 settings after a test-size guardrail; one flip is bootstrap-supported. Thus p-core temporal benchmarks require an explicit filter/split order, not just a pruning threshold. We recommend reporting fitted-train core violation and publishing split manifests as lightweight checks for reproducible temporal evaluation.
Last synced
This work has 1 recorded citation, but citing papers have not been linked locally yet.
No comments yet — start the discussion below.