Individual and group fairness assessments via counterfactual explanations
Abstract This study explores the potential of counterfactual explanations to assess artificial intelligence (AI) fairness, especially in critical decision-making systems. Predictive models may amplify biases inherent in data sets or algorithms, and given the absence of a universally accepted fairness metric, a case-specific approach becomes mandatory. Existing statistical fairness metrics may not capture all aspects that are relevant to a context-aware assessment of non-discrimination. The goal of this work is to define a measure of fairness for AI systems based on explainable artificial intelligence concepts. Specifically, it leverages the analysis of counterfactual explanations of individuals/groups and their comparison with similar individuals/groups. Compared to existing state-of-the-art works, the contributions are (i) extending the definition of individual fairness, not limiting unfairness to decisions based on sensitive attributes but also ensuring similar treatment amongst similar individuals; (ii) revisiting (and generalising) existing notions and introducing new, more refined notions of group fairness based on counterfactuals; (iii) defining quantitative fairness metrics that reflect the evidence gathered through the analysis/comparison of counterfactual explanations.
Authors
- Federico Sabbatini (ORCID: https://orcid.org/0000-0002-0532-6777)
- Roberta Calegari (ORCID: https://orcid.org/0000-0003-3794-2942)
Institutions
- University of Urbino (IT)
- University of Bologna (IT)
Publication Details
- Journal
- AI and Ethics
- Published
- 2026-09-30
- DOI
- https://doi.org/10.1007/s43681-026-01411-w
- Primary Topic
- Ethics and Social Impacts of AI
- Type
- article
- Field-Weighted Citation Impact
- 0.00