Evaluation methods and metrics towards trustworthy AI
Abstract The ongoing developments in Artificial Intelligence (AI) are introducing new technologies, initiatives, and applications. To assure that these developments do not translate into potential risks that are not captured, understood, measured, and managed, a trustworthiness approach is necessary. AI trustworthiness is treated as a multidimensional construct that spans transparency, safety, security, privacy, robustness, and fairness. We argue that these properties they should be translated into explicit, testable criteria. This article aims to synthesize evaluation approaches that make AI trustworthiness observable and measurable, enabling stakeholders to identify gaps, prioritize improvements, and align system behaviour with ethical, social, and legal values. A scoping review (resulting in 27 papers) is augmented with knowledge of 6 experts to map (i) the main trustworthy AI aspects, (ii) the evaluation methods used to assess these aspects, (iii) the metrics applied within them, and (iv) recommendations proposed to strengthen trustworthiness. The results suggest that current evaluation practices interpret and operationalize trustworthy AI in diverse ways. They often rely on established methods or benchmarks that may imply bias, limited comparability, or provide limited insight into the reasoning quality. We recommend to design evaluation frameworks that explicitly target trustworthiness aspects from a multidisciplinary angle and being instantiated to specific contexts or problems.
Authors
- Catharina Margaretha van Leersum (ORCID: https://orcid.org/0000-0002-1003-0794)
- Clara Maathuis (ORCID: https://orcid.org/0000-0003-3483-1569)
Institutions
- Open University of the Netherlands (NL)
Publication Details
- Journal
- AI and Ethics
- Published
- 2026-09-18
- DOI
- https://doi.org/10.1007/s43681-026-01349-z
- Primary Topic
- Ethics and Social Impacts of AI
- Type
- article
- Field-Weighted Citation Impact
- 0.00