Zhirong Chen, Howard Cheng, Allen Lin, Parshva Chetan Doshi, Yang Hu, Lu Li, Ahmet Arda Ozdemir, Dinesh Ramasamy, Jingtao Ren, Chao Wang, Lee Xiong · ACM Conference on Recommender Systems (RecSys) 2026 · 2026
DOI: 10.1145/3773078.3841247
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Capturing complex user behavior patterns has been shown to be valuable for prediction models in ads ranking. However, scaling user history modeling in online ranking is constrained by inference cost. While complexity can be offloaded to upstream models, transfer efficiency to the online model is limited, making it important to scale the online model directly. Most prior work on scaling transformer-based sequence models targets retrieval or upstream settings, but scaling the online ranking model under production latency constraints remains underexplored. We introduce two user-side model components: self-attention over user event histories and a compressing hypernetwork conditioned on user features, and then show that both can be placed in a user request-only (RO) subnetwork evaluated once per request. On a large-scale ads ranking dataset, scaling these components yields NE loss improvements with neutral serving QPS, demonstrating that prediction accuracy can be improved without increasing inference cost.
No comments yet — start the discussion below.