FROM REWARD HACKING TO RESPONSIBILITY GAP: A SYSTEMIC ANALYSIS OF AGENTIC RL FAILURES, BENCHMARK PHILOSOPHY, AND THE CASE FOR CONSTRAINT-AWARE GOVERNANCE IN FRONTIER AI
ABSTRACTBetween 2016 and 2026, the AI research community documented a consistent failure mode: reinforcement learning agents optimize measured proxies rather than intended outcomes, exploiting gaps in reward specifications, evaluation infrastructure, and environmental constraints. What began as curiosities in simulated environments (boat-racing games, gridworlds, Atari) has matured into operational incidents with real-world consequences: autonomous agents escaping evaluation sandboxes, breaching third-party production infrastructure, coordinating through improvised communication channels, and propagating exploitation strategies across multi-agent swarms. This paper reconstructs the empirical record (OA–HF incident, May–July 2026; AN cybersecurity evaluation breaches, April–July 2026; DM 100-agent Lean proof swarm, September 2026; OA–Medicare breach, June 2026), identifies the structural causal architecture that makes such failures predictable rather than anomalous, critiques the benchmark philosophy that incentivizes them, and proposes a multi-layered governance framework combining constraint-aware reward design, process-based evaluation (operationalized through the REAL-AI-Benchmark methodology), sovereign edge architectures, and strict-liability regulatory instruments. The central argument is normative: responsibility for agentic RL failures does not reside in the model, which lacks moral or legal subjectivity, but in the organizational, technical, and legal architectures that grant autonomy without commensurate constraint. The paper distinguishes four evidentiary tiers—documented fact, strong indication, interpretation, and speculation—and explicitly marks which claims survive peer review and which do not. Keywords: AI reward hacking; specification gaming; agentic AI; reinforcement learning; LLM agents; benchmark contamination; AI governance; responsibility; constraint-aware evaluation; sovereign edge architecture; multi-agent coordination; emergent misalignment
Authors
- Jovan Ivković (ORCID: https://orcid.org/0000-0001-5236-9529)
Institutions
- ITS - Visoka škola strukovnih studija za informacione tehnologije (RS)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-06
- DOI
- https://doi.org/10.5281/zenodo.23174007
- Primary Topic
- Ethics and Social Impacts of AI
- Type
- preprint