Alessandro Petruzzelli, Alessandro Francesco Maria Martina, Cataldo Musto, Marco de Gemmis, Pasquale Lops, Giovanni Maria Semeraro · ACM Conference on Recommender Systems (RecSys) 2026 · 2026
DOI: 10.1145/3773078.3831840
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Agentic Conversational Recommender Systems (ACRSs) are designed to recommend through multi-turn dialogue with users whose needs are not fully formed at the outset. However, their evaluation almost exclusively relies on user simulators that instantiate users with clear, pre-formed needs, reducing the interaction to a retrieval over attributes disclosed in the initial turns of the conversation. This covers only a narrow slice of the behaviors real users exhibit, and assessing the robustness and reliability of these systems requires simulated users that span a wider range. To this end, we introduce a stereotype-conditioned, open-weight user simulator that spans three behavioral stereotypes: \textit{Direct}, \textit{Vague-Proactive}, and \textit{Vague-Reactive}. Benchmarking four state-of-the-art ACRSs across four e-commerce domains with our simulator, three findings emerge. First, under certain stereotypes, the user stops contributing new information about the target as turns accumulate. At the same time, the agent continues to act, a regime previously unobserved, which we name \emph{Unproductive Stagnation} and formalize via Preference Coverage. Second, a systematic \emph{Robustness Gap} emerges: as the simulated user shifts from decisive to passive, accuracy collapses while conversations grow longer. Third, accuracy degrades more sharply than Preference Coverage does, decomposing the gap into two separable capabilities current ACRSs lack: elicitation and retrieval, which current evaluation entangles in a single score. Our simulator makes this distinction reportable and gives the field a controllable axis along which elicitation and retrieval can be measured and compared.
No comments yet — start the discussion below.