Emrul Hasan, Sajib Saha, Chen Ding, Jimmy Xiangji Huang, John-Jose Nuñez, Shaina Raza · Preprints.org 2026 · 2026
DOI: 10.20944/preprints202609.1665.v1
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Reproducibility underpins scientific progress, yet in recommendation systems (RS) research, it remains one of the most persistent and underappreciated problems. Algorithms have advanced rapidly, but reported performance gains are often fragile and difficult to verify independently. To address this gap, using the PRISMA framework, we systematically review peer-reviewed studies published between 2013, when reproducibility began receiving attention in the RS community, and 2026. We analyze all identified papers and introduce a six-category hierarchical taxonomy of reproducibility challenges in RS. Building on this, we synthesize the solutions proposed in prior work into a structured framework that maps each category of solutions to the challenges it addresses. Our review shows that reproducibility in RS is a property of the entire experimental workflow, not simply the availability of code and data. To this end, we propose an end-to-end Reproducibility Pipeline, spanning data construction, splitting, candidate generation, configuration, training or execution, evaluation, and reporting. Recent advances in LLM-based and agentic recommendation systems introduce new reproducibility risks, including prompt sensitivity, nondeterministic inference, proprietary APIs, multi-step reasoning, tool use, and external dependencies. Finally, we outline future directions for reproducibility in RS. Relevant materials are available at Github.
No comments yet — start the discussion below.