Federico Sabbatini, Roberta Calegari · AI and Ethics 2026 · 2026
DOI: 10.1007/s43681-026-01411-w
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
This study explores the potential of counterfactual explanations to assess artificial intelligence (AI) fairness, especially in critical decision-making systems. Predictive models may amplify biases inherent in data sets or algorithms, and given the absence of a universally accepted fairness metric, a case-specific approach becomes mandatory. Existing statistical fairness metrics may not capture all aspects that are relevant to a context-aware assessment of non-discrimination. The goal of this work is to define a measure of fairness for AI systems based on explainable artificial intelligence concepts. Specifically, it leverages the analysis of counterfactual explanations of individuals/groups and their comparison with similar individuals/groups. Compared to existing state-of-the-art works, the contributions are (i) extending the definition of individual fairness, not limiting unfairness to decisions based on sensitive attributes but also ensuring similar treatment amongst similar individuals; (ii) revisiting (and generalising) existing notions and introducing new, more refined notions of group fairness based on counterfactuals; (iii) defining quantitative fairness metrics that reflect the evidence gathered through the analysis/comparison of counterfactual explanations.
No comments yet — start the discussion below.