Pre-Irreversibility Safety for Advanced AI: Distributed Irreversible Authority, Timely Refusal, and Human-Revisable Development
Advanced-AI safety is commonly approached through alignment, capability evaluation, control, monitoring, human oversight, governance, and systemic-risk analysis. Existing work also explicitly treats irreversibility, capability-authority separation, meaningful human control, gradual human disempowerment, and distributed trust. The unresolved problem addressed here is therefore not the introduction of irreversibility as a safety concern, but its operational decomposition into distinct, jointly required evaluation objects. This paper proposes a pre-irreversibility safety framework organized around three non-substitutable candidate dimensions. Distributed Irreversible Authority requires that no sufficiently small coalition of independently compromisable human or artificial authority domains can force a civilizationally irreversible transition. Effective Pre-Irreversibility Refusal requires that a dangerous transition remain practically interruptible before the last safe intervention point; nominal shutdown authority is insufficient when detection, authorization, enforcement, or verification take too long. Human-Revisable Development Viability requires that ordinary technological and institutional development remain within a region from which human-governed procedures retain realistic pathways to materially different futures. The framework associates these dimensions with three formal objects: a robust existential-authority threshold κE, a robust refusal margin M_R^rob, and a finite-horizon human-revisable viability kernel K_R^H. Counterexamples show that the three dimensions are non-substitutable: satisfying any two does not in general imply the third. Boundary Completeness and Epistemic Independence are treated separately as assurance conditions governing whether claims about the three substantive dimensions are credible. Three worked toy reference cases instantiate the formal objects: a threshold-release architecture with hidden shared-root dependence, a time-bounded emergency-refusal scenario, and a cumulative fallback-capacity model. The cases establish that the definitions are executable, the evaluation procedure is reproducible under declared assumptions, and parameter boundaries can be independently recomputed. They do not constitute empirical validation of real frontier-AI deployments. The paper does not claim that the three dimensions are sufficient for safe superintelligence. Its narrower candidate contribution is an operational integration of Authority, Refusal, and Revisability into a single pre-irreversibility safety case that can be inspected, calculated, challenged, and revised.
Authors
- Elias Arden
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-30
- DOI
- https://doi.org/10.5281/zenodo.23042906
- Primary Topic
- Ethics and Social Impacts of AI
- Type
- preprint