Real-world AI evaluation design and planning

Understanding how AI systems behave in the real- world is becoming more imperative in a world where companies, organizations, and governments are quickly adopting and deploying this technology. Using a novel framework for real-world AI evaluation, CIRCLE [1], we present a set of activities for testing AI systems in deployment contexts including field testing and red teaming. We demonstrate how these activities can produce specific outcomes of interest to stakeholders outside the AI stack. The CIRCLE framework is rooted in an understanding of the AI lifecycle that moves beyond traditional model-centric evaluation techniques. By providing a hypothetical case study from an education setting, we showcase how evaluation approaches that are responsive to stakeholders’ views outside of the traditional AI stack allow for systems that are aligned with stakeholder objectives, support the aims of building more trustworthy and safer AI systems, and enable better decisions about their deployment.

Authors

Publication Details

Journal
Bournemouth University Research Online (Bournemouth University)
Published
2026-06-12
Primary Topic
Ethics and Social Impacts of AI
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Real-world AI evaluation design and planning

M. Briggs, Carina Westling, T. Skeadas
Bournemouth University Research Online (Bournemouth University)
Ethics and Social Impacts of AI
article

Real-world AI evaluation design and planning

M. Briggs, Carina Westling, T. Skeadas
article en

Abstract

Understanding how AI systems behave in the real- world is becoming more imperative in a world where companies, organizations, and governments are quickly adopting and deploying this technology. Using a novel framework for real-world AI evaluation, CIRCLE [1], we present a set of activities for testing AI systems in deployment contexts including field testing and red teaming. We demonstrate how these activities can produce specific outcomes of interest to stakeholders outside the AI stack. The CIRCLE framework is rooted in an understanding of the AI lifecycle that moves beyond traditional model-centric evaluation techniques. By providing a hypothetical case study from an education setting, we showcase how evaluation approaches that are responsive to stakeholders’ views outside of the traditional AI stack allow for systems that are aligned with stakeholder objectives, support the aims of building more trustworthy and safer AI systems, and enable better decisions about their deployment.

Bournemouth University Research Online (Bournemouth University)
Openalex Percentile: Top 25%
Ethics and Social Impacts of AI
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.