AccessAgent-Bench: Evaluating Accessibility Preservation in Long-Horizon Web Development
AI coding agents can implement features that satisfy functional tests while introducing barriers for people who use keyboards, screen readers, or magnification. Evaluating a final application alone does not reveal when such barriers emerge, whether they persist, or whether later changes repair them. This paper proposes AccessAgent-Bench, a benchmark for evaluating accessibility preservation across cumulative web-development tasks. The proposed suite comprises three application domains, each with eight sequential feature requests, evaluated under accessibility-agnostic and accessibility-aware instructions. Agents retain their own evolving code across stages. Each checkpoint combines functional testing, automated accessibility analysis, and structured expert review. We define Accessibility Preservation Rate over stable, previously passing checks and complement it with baseline survival, new-feature accessibility, defect persistence, and functional completion. Explicit treatment of blocked evaluations and changing applicability prevents missing features or inaccessible states from being counted as successes. A blocked, repeated-run study will estimate instruction effects while reporting cost and completion time separately. The contribution is a prospective evaluation protocol connecting cumulative software evolution to accessibility regression measurement. No agent experiments, validated benchmark repositories, or expert annotations are reported in this version. The protocol specifies the work required to obtain those results reproducibly and to distinguish evidence of preservation from unsupported claims of full accessibility conformance.
Authors
- Surafel Workayehu Terefe
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-28
- DOI
- https://doi.org/10.5281/zenodo.23012265
- Primary Topic
- Digital Accessibility for Disabilities
- Type
- preprint