Amir Reza Mohammadi, Andreas Peintner, Michael Müller, Eva Zangerle · ACM Transactions on Recommender Systems 2026 · 2026
DOI: 10.1145/3839564
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Explainability in recommender systems (RS) remains a pivotal challenge. Counterfactual explanations have emerged as a particularly actionable paradigm, offering intuitive “what-if” reasoning. However, their evaluation lacks principled standards. Current metrics primarily assess whether explanations change the top-ranked recommendation, overlooking two fundamental aspects of explanation quality. First, evaluation results are frequently inconsistent , as metrics are tightly coupled to the underlying recommender’s performance. Second, explanations are rarely assessed for compactness —whether changes are sufficiently small to remain interpretable. Oversized counterfactuals may technically succeed but fail to provide practical insight. In this work, we advocate for a holistic evaluation perspective centered on consistency and compactness . We systematically analyze how extending evaluation beyond top-1 to top-k recommendations improves metric stability and reduces dependence on recommender fluctuations. In parallel, we introduce compactness-aware evaluation criteria that quantify the minimality of counterfactual modifications. Through extensive experiments across multiple datasets and models, we demonstrate that jointly considering these dimensions yields more reliable assessments. Our findings expose key factors driving metric instability, highlight the trade-off between effectiveness and explanation size, and provide practical guidelines toward standardized, faithful, and compact evaluation of counterfactual explanations in recommender systems.
No comments yet — start the discussion below.