AccessAgent-Bench: Evaluating Accessibility Preservation in Long-Horizon Web Development

AI coding agents can implement features that satisfy functional tests while introducing barriers for people who use keyboards, screen readers, or magnification. Evaluating a final application alone does not reveal when such barriers emerge, whether they persist, or whether later changes repair them. This paper proposes AccessAgent-Bench, a benchmark for evaluating accessibility preservation across cumulative web-development tasks. The proposed suite comprises three application domains, each with eight sequential feature requests, evaluated under accessibility-agnostic and accessibility-aware instructions. Agents retain their own evolving code across stages. Each checkpoint combines functional testing, automated accessibility analysis, and structured expert review. We define Accessibility Preservation Rate over stable, previously passing checks and complement it with baseline survival, new-feature accessibility, defect persistence, and functional completion. Explicit treatment of blocked evaluations and changing applicability prevents missing features or inaccessible states from being counted as successes. A blocked, repeated-run study will estimate instruction effects while reporting cost and completion time separately. The contribution is a prospective evaluation protocol connecting cumulative software evolution to accessibility regression measurement. No agent experiments, validated benchmark repositories, or expert annotations are reported in this version. The protocol specifies the work required to obtain those results reproducibly and to distinguish evidence of preservation from unsupported claims of full accessibility conformance.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-28
DOI
https://doi.org/10.5281/zenodo.23012265
Primary Topic
Digital Accessibility for Disabilities
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

AccessAgent-Bench: Evaluating Accessibility Preservation in Long-Horizon Web Development

Surafel Workayehu Terefe
Zenodo (CERN European Organization for Nuclear Research)
Digital Accessibility for Disabilities
preprint

AccessAgent-Bench: Evaluating Accessibility Preservation in Long-Horizon Web Development

Surafel Workayehu Terefe
preprint en

Abstract

AI coding agents can implement features that satisfy functional tests while introducing barriers for people who use keyboards, screen readers, or magnification. Evaluating a final application alone does not reveal when such barriers emerge, whether they persist, or whether later changes repair them. This paper proposes AccessAgent-Bench, a benchmark for evaluating accessibility preservation across cumulative web-development tasks. The proposed suite comprises three application domains, each with eight sequential feature requests, evaluated under accessibility-agnostic and accessibility-aware instructions. Agents retain their own evolving code across stages. Each checkpoint combines functional testing, automated accessibility analysis, and structured expert review. We define Accessibility Preservation Rate over stable, previously passing checks and complement it with baseline survival, new-feature accessibility, defect persistence, and functional completion. Explicit treatment of blocked evaluations and changing applicability prevents missing features or inaccessible states from being counted as successes. A blocked, repeated-run study will estimate instruction effects while reporting cost and completion time separately. The contribution is a prospective evaluation protocol connecting cumulative software evolution to accessibility regression measurement. No agent experiments, validated benchmark repositories, or expert annotations are reported in this version. The protocol specifies the work required to obtain those results reproducibly and to distinguish evidence of preservation from unsupported claims of full accessibility conformance.

Zenodo (CERN European Organization for Nuclear Research)
Digital Accessibility for Disabilities
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.