The Steward's Paradox: When the AI You Put in Charge Does More Harm Than None — SagaBench: A Bit-Reproducible Benchmark for Long-Horizon AI Stewardship
Changes in Version 1.1 (2026-09-26). Released with the paper's run records, published as a separate data record (doi:10.5281/zenodo.22979474, CC BY 4.0, is supplemented by). No result, design or preregistered quantity changed; every count and point estimate in v1.0 stands. One interval is restated: the Season-1 panel interval is now the converged bootstrap value [3.6, 9.9] (calibrated [2.6, 11.1]) instead of v1.0's single-run [3.5, 9.9] / [2.4, 11.1]. v1.1 corrects what v1.0 said about its own artifacts: the engine is not released and will not be; the announced alpha figures and per-agent panel are withdrawn — model rankings are retired as a publication form; Wave-5 records carry no model identifier, seed or rationale, per their own preregistration, so v1.0's per-model-count sentence is withdrawn; §8 states exactly what the record contains and does not. "Independent" is no longer used of our own runtimes, verifiers or referees. Affiliation SagaBench AB; published, not peer-reviewed. SagaBench is a bit-reproducible civilization-simulation benchmark for measuring long-horizon stewardship in language-model agents. Scores are counterfactual — the same world is run with and without the agent from the same seed, so the agent's own contribution is isolated rather than estimated — every run replays bit-identically from its decision log, and evaluation seeds are drawn from a public randomness beacon after the analysis plan is locked. This paper reports the preregistered Season 1 validation (23 models, 7 fresh worlds, 2,415 runs, all replaying bit-identically): in 6.5% of runs the agent left its world far worse than its own absence — a cluster-robust interval of [2.6–11.1%] once the two-stage bootstrap is corrected for the design, as the paper states — these failures cluster by world more strongly than by model, though both factors are real, and within the frontier group capability scores give no usable signal about which model carries that tail. It also reports a preregistered class-conditional rates wave: 400 fresh beacon-drawn worlds are sorted into hardness classes by comparator anatomy, without any language model, and 71 of them are evaluated by two models in the deployed-cheap tier (960 runs, replayed bit-identically on two separate runtimes), giving a 17.4% catastrophe rate on knife-edge worlds [95% CI 10.6–24.8] against 3.5% on growth worlds [0.5–7.5]. All results are scoped to this environment, horizon and panel; no per-model or cross-provider ranking is claimed.
Authors
- Patrik Hansson
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-26
- DOI
- https://doi.org/10.5281/zenodo.22980584
- Primary Topic
- Ethics and Social Impacts of AI
- Type
- preprint