Maxim Skurikhin, Kiryl Liakhnovich, Oleg Lashinin · ACM Conference on Recommender Systems (RecSys) 2026 · 2026
DOI: 10.1145/3773078.3831815
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
LLM-enhanced linear autoencoders (L3AE) incorporate semantic item representations from large language models into collaborative filtering and demonstrate significant gains on long-tail items. However, L3AE requires dense n × n matrices, which makes it impractical for large catalogs — exactly where semantic enrichment would help most. We propose scalable formulations that factor the semantic similarity matrix into a compact low-rank form and compute all required precision-matrix products without ever building any dense n × n matrix. Our method preserves the algebraic structure of L3AE while replacing its dense collaborative precision with SANSA’s sparse approximate inverse, enabling efficient scaling to industrial-size catalogs. Experiments on several public datasets show consistent gains over multiple baselines, and these gains tend to grow in colder, sparser catalogs — a pattern consistent with semantic enrichment helping most where collaborative signals are weakest. Code and configs are available at https://github.com/ViV99/scal3ae.
No comments yet — start the discussion below.