Shuli Wang, Junwei Yin, Changhao Li, S. C. Kou, Chi Wang, Yinqiu Huang, Yinhua Zhu, Haitao Wang, Xingxing Wang · · 2026
DOI: 10.1145/3773078.3831746
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Scaling industrial recommender models has followed two parallel paradigms: sample information scaling—enriching the information content of each training sample through deeper and longer behavior sequences—and model capacity scaling—unifying sequence modeling and feature interaction within a single Transformer backbone. However, these two paradigms still face two structural limitations. Firstly, sample information scaling methods encode only a subset of each historical interaction into the sequence token, leaving the majority of the original sample context unexploited and precluding the modeling of sample-level, time-varying features. Secondly, model capacity scaling methods are inherently constrained by the structural heterogeneity between sequential and non-sequential features, preventing the model from fully realizing its representational capacity.
No comments yet — start the discussion below.