Pre-Irreversibility Safety for Advanced AI: Distributed Irreversible Authority, Timely Refusal, and Human-Revisable Development

Advanced-AI safety is commonly approached through alignment, capability evaluation, control, monitoring, human oversight, governance, and systemic-risk analysis. Existing work also explicitly treats irreversibility, capability-authority separation, meaningful human control, gradual human disempowerment, and distributed trust. The unresolved problem addressed here is therefore not the introduction of irreversibility as a safety concern, but its operational decomposition into distinct, jointly required evaluation objects. This paper proposes a pre-irreversibility safety framework organized around three non-substitutable candidate dimensions. Distributed Irreversible Authority requires that no sufficiently small coalition of independently compromisable human or artificial authority domains can force a civilizationally irreversible transition. Effective Pre-Irreversibility Refusal requires that a dangerous transition remain practically interruptible before the last safe intervention point; nominal shutdown authority is insufficient when detection, authorization, enforcement, or verification take too long. Human-Revisable Development Viability requires that ordinary technological and institutional development remain within a region from which human-governed procedures retain realistic pathways to materially different futures. The framework associates these dimensions with three formal objects: a robust existential-authority threshold κE, a robust refusal margin M_R^rob, and a finite-horizon human-revisable viability kernel K_R^H. Counterexamples show that the three dimensions are non-substitutable: satisfying any two does not in general imply the third. Boundary Completeness and Epistemic Independence are treated separately as assurance conditions governing whether claims about the three substantive dimensions are credible. Three worked toy reference cases instantiate the formal objects: a threshold-release architecture with hidden shared-root dependence, a time-bounded emergency-refusal scenario, and a cumulative fallback-capacity model. The cases establish that the definitions are executable, the evaluation procedure is reproducible under declared assumptions, and parameter boundaries can be independently recomputed. They do not constitute empirical validation of real frontier-AI deployments. The paper does not claim that the three dimensions are sufficient for safe superintelligence. Its narrower candidate contribution is an operational integration of Authority, Refusal, and Revisability into a single pre-irreversibility safety case that can be inspected, calculated, challenged, and revised.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-30
DOI
https://doi.org/10.5281/zenodo.23042905
Primary Topic
Ethics and Social Impacts of AI
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

Pre-Irreversibility Safety for Advanced AI: Distributed Irreversible Authority, Timely Refusal, and Human-Revisable Development

Elias Arden
Zenodo (CERN European Organization for Nuclear Research)
Ethics and Social Impacts of AI
preprint

Pre-Irreversibility Safety for Advanced AI: Distributed Irreversible Authority, Timely Refusal, and Human-Revisable Development

Elias Arden
preprint en

Abstract

Advanced-AI safety is commonly approached through alignment, capability evaluation, control, monitoring, human oversight, governance, and systemic-risk analysis. Existing work also explicitly treats irreversibility, capability-authority separation, meaningful human control, gradual human disempowerment, and distributed trust. The unresolved problem addressed here is therefore not the introduction of irreversibility as a safety concern, but its operational decomposition into distinct, jointly required evaluation objects. This paper proposes a pre-irreversibility safety framework organized around three non-substitutable candidate dimensions. Distributed Irreversible Authority requires that no sufficiently small coalition of independently compromisable human or artificial authority domains can force a civilizationally irreversible transition. Effective Pre-Irreversibility Refusal requires that a dangerous transition remain practically interruptible before the last safe intervention point; nominal shutdown authority is insufficient when detection, authorization, enforcement, or verification take too long. Human-Revisable Development Viability requires that ordinary technological and institutional development remain within a region from which human-governed procedures retain realistic pathways to materially different futures. The framework associates these dimensions with three formal objects: a robust existential-authority threshold κE, a robust refusal margin M_R^rob, and a finite-horizon human-revisable viability kernel K_R^H. Counterexamples show that the three dimensions are non-substitutable: satisfying any two does not in general imply the third. Boundary Completeness and Epistemic Independence are treated separately as assurance conditions governing whether claims about the three substantive dimensions are credible. Three worked toy reference cases instantiate the formal objects: a threshold-release architecture with hidden shared-root dependence, a time-bounded emergency-refusal scenario, and a cumulative fallback-capacity model. The cases establish that the definitions are executable, the evaluation procedure is reproducible under declared assumptions, and parameter boundaries can be independently recomputed. They do not constitute empirical validation of real frontier-AI deployments. The paper does not claim that the three dimensions are sufficient for safe superintelligence. Its narrower candidate contribution is an operational integration of Authority, Refusal, and Revisability into a single pre-irreversibility safety case that can be inspected, calculated, challenged, and revised.

Zenodo (CERN European Organization for Nuclear Research)
Peace, Justice and strong institutions
Ethics and Social Impacts of AI
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.