Jinfeng Xu, Zheyu Chen, Ziyue Peng, Wei Wang, Xiping Hu, Edith C. H. Ngai, Victor C. M. Leung · Preprints.org 2026 · 2026
DOI: 10.20944/preprints202609.1653.v1
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Evaluation validity can change which recommender appears better. Ranking a relevant item supplied by an evaluator differs from finding it in a catalog; generating a valid identifier differs from identifying one item. This focused survey examines 20 offline protocol studies within a 49-paper coded corpus and develops two findings. First, reported model or history-strategy orderings change when studies vary candidate pools, credit for ambiguous identifiers, or the benefit measured by fairness. Second, apparently conflicting results can concern different objects: selecting long-tail recommendations differs from recovering popular benchmark records, and reducing attribute recoverability differs from distributing user benefit. We explain these findings through the items a model can access, the outputs receiving credit, the held-out population, and the outcome being measured. The synthesis derives evaluation controls for interpreting utility, diversity, exposure, and fairness in LLM-based and generative recommendation.
No comments yet — start the discussion below.