Anna Volodkevich, Anton Klenitskiy, Artem Fatkulin, Daria Denisova, Alexey Vasilev · ACM Conference on Recommender Systems (RecSys) 2026 · 2026
DOI: 10.1145/3773078.3841287
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Generative recommenders represent each item as a Semantic ID: a short sequence of discrete tokens. Popular quantization algorithms, such as RQ-VAE and residual K-means, can map several items to the same token sequence, producing collisions. Collisions must be resolved before training a single-stage generative recommender, since the model cannot distinguish items that share an identifier. The widely adopted disambiguation technique appends one extra codebook and uses the last token as an arbitrary item counter within each collision group. We argue that this last codebook can be used more effectively as a source of additional content or collaborative signal. Keeping the quantizer, the generative model, and the Semantic ID length unchanged, we replace the ordinal counter with a token that both resolves collisions and carries a useful signal. Across several datasets, the collision resolution choice alone could yield an easy-to-implement, quantizer-agnostic gain in recommendation quality (up to +12.7% in NDCG@10). The code is available at https://github.com/sb-ai-lab/sid-collisions/tree/recsys_2026.
No comments yet — start the discussion below.