The Steward's Paradox: When the AI You Put in Charge Does More Harm Than None — SagaBench: A Bit-Reproducible Benchmark for Long-Horizon AI Stewardship

Changes in Version 1.1 (2026-09-26). Released with the paper's run records, published as a separate data record (doi:10.5281/zenodo.22979474, CC BY 4.0, is supplemented by). No result, design or preregistered quantity changed; every count and point estimate in v1.0 stands. One interval is restated: the Season-1 panel interval is now the converged bootstrap value [3.6, 9.9] (calibrated [2.6, 11.1]) instead of v1.0's single-run [3.5, 9.9] / [2.4, 11.1]. v1.1 corrects what v1.0 said about its own artifacts: the engine is not released and will not be; the announced alpha figures and per-agent panel are withdrawn — model rankings are retired as a publication form; Wave-5 records carry no model identifier, seed or rationale, per their own preregistration, so v1.0's per-model-count sentence is withdrawn; §8 states exactly what the record contains and does not. "Independent" is no longer used of our own runtimes, verifiers or referees. Affiliation SagaBench AB; published, not peer-reviewed. SagaBench is a bit-reproducible civilization-simulation benchmark for measuring long-horizon stewardship in language-model agents. Scores are counterfactual — the same world is run with and without the agent from the same seed, so the agent's own contribution is isolated rather than estimated — every run replays bit-identically from its decision log, and evaluation seeds are drawn from a public randomness beacon after the analysis plan is locked. This paper reports the preregistered Season 1 validation (23 models, 7 fresh worlds, 2,415 runs, all replaying bit-identically): in 6.5% of runs the agent left its world far worse than its own absence — a cluster-robust interval of [2.6–11.1%] once the two-stage bootstrap is corrected for the design, as the paper states — these failures cluster by world more strongly than by model, though both factors are real, and within the frontier group capability scores give no usable signal about which model carries that tail. It also reports a preregistered class-conditional rates wave: 400 fresh beacon-drawn worlds are sorted into hardness classes by comparator anatomy, without any language model, and 71 of them are evaluated by two models in the deployed-cheap tier (960 runs, replayed bit-identically on two separate runtimes), giving a 17.4% catastrophe rate on knife-edge worlds [95% CI 10.6–24.8] against 3.5% on growth worlds [0.5–7.5]. All results are scoped to this environment, horizon and panel; no per-model or cross-provider ranking is claimed.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-26
DOI
https://doi.org/10.5281/zenodo.22980584
Primary Topic
Ethics and Social Impacts of AI
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

The Steward's Paradox: When the AI You Put in Charge Does More Harm Than None — SagaBench: A Bit-Reproducible Benchmark for Long-Horizon AI Stewardship

Patrik Hansson
Zenodo (CERN European Organization for Nuclear Research)
Ethics and Social Impacts of AI
preprint

The Steward's Paradox: When the AI You Put in Charge Does More Harm Than None — SagaBench: A Bit-Reproducible Benchmark for Long-Horizon AI Stewardship

Patrik Hansson
preprint en

Abstract

Changes in Version 1.1 (2026-09-26). Released with the paper's run records, published as a separate data record (doi:10.5281/zenodo.22979474, CC BY 4.0, is supplemented by). No result, design or preregistered quantity changed; every count and point estimate in v1.0 stands. One interval is restated: the Season-1 panel interval is now the converged bootstrap value [3.6, 9.9] (calibrated [2.6, 11.1]) instead of v1.0's single-run [3.5, 9.9] / [2.4, 11.1]. v1.1 corrects what v1.0 said about its own artifacts: the engine is not released and will not be; the announced alpha figures and per-agent panel are withdrawn — model rankings are retired as a publication form; Wave-5 records carry no model identifier, seed or rationale, per their own preregistration, so v1.0's per-model-count sentence is withdrawn; §8 states exactly what the record contains and does not. "Independent" is no longer used of our own runtimes, verifiers or referees. Affiliation SagaBench AB; published, not peer-reviewed. SagaBench is a bit-reproducible civilization-simulation benchmark for measuring long-horizon stewardship in language-model agents. Scores are counterfactual — the same world is run with and without the agent from the same seed, so the agent's own contribution is isolated rather than estimated — every run replays bit-identically from its decision log, and evaluation seeds are drawn from a public randomness beacon after the analysis plan is locked. This paper reports the preregistered Season 1 validation (23 models, 7 fresh worlds, 2,415 runs, all replaying bit-identically): in 6.5% of runs the agent left its world far worse than its own absence — a cluster-robust interval of [2.6–11.1%] once the two-stage bootstrap is corrected for the design, as the paper states — these failures cluster by world more strongly than by model, though both factors are real, and within the frontier group capability scores give no usable signal about which model carries that tail. It also reports a preregistered class-conditional rates wave: 400 fresh beacon-drawn worlds are sorted into hardness classes by comparator anatomy, without any language model, and 71 of them are evaluated by two models in the deployed-cheap tier (960 runs, replayed bit-identically on two separate runtimes), giving a 17.4% catastrophe rate on knife-edge worlds [95% CI 10.6–24.8] against 3.5% on growth worlds [0.5–7.5]. All results are scoped to this environment, horizon and panel; no per-model or cross-provider ranking is claimed.

Zenodo (CERN European Organization for Nuclear Research)
Peace, Justice and strong institutions
Ethics and Social Impacts of AI
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.