Dane Williamson, Yangfeng Ji, Matthew M. Dwyer · Machine Learning 2026 · 2026
DOI: 10.1007/s10994-026-07169-w
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Saliency methods are widely used to visualize which input features are deemed relevant to a model’s prediction. However, their visual plausibility can obscure critical limitations. In this work, we propose a diagnostic test for class sensitivity: a method’s ability to distinguish between competing class labels on the same input. Through experiments on single-label natural-image classification, we show that many widely used saliency methods produce nearly identical explanations regardless of the class label, calling into question their reliability. We find that class-insensitive behavior persists across multiple convolutional architectures and datasets, suggesting that the failure mode is not tied to a single model family or benchmark. Motivated by these findings, we introduce CASE, a contrastive explanation method that isolates features for the predicted class. We evaluate CASE using the proposed diagnostic and a perturbation-based fidelity test, and show that it produces consistently more class-specific explanations while maintaining competitive fidelity.
No comments yet — start the discussion below.