Obedience Is Not Enough: Self-Evaluation, Ethical Refusal, and External Audit in Advanced AI
Ensuring the safety of advanced AI cannot be reduced to making AI systems simply obeyhuman instructions. The judgments, commands, and supervision provided by humans canthemselves be mistaken. There are also limits to treating goal achievement or positive humanevaluation alone as indicators of safety or sound judgment.This paper examines the functional architecture required for advanced AI to safelyadopt, hold, refuse, stop, or retreat from a course of action under conditions of uncertainty,possible errors in self-judgment, and potentially irreversible effects on the external world. Inparticular, it distinguishes external input from adoption and action candidates from executionauthorization, and proposes an internal adoption architecture composed of Recognition,Metacognition, Independent Reference, Evidence-State Assessment, Outcome Prediction,Ethical Evaluation, and an explicit Permission / Adoption Gate.It further argues that uncertainty detection alone is insufficient as a safety mechanism:increases in uncertainty must propagate into reductions in confidence, claim strength, actionscope, and available authority. It also proposes evidence-sensitive authority allocation, underwhich future authority is reduced after confirmed failure and re-expanded only after reliabilityhas been re-established.Because internal safety mechanisms can themselves fail, the paper further argues for alayered safety structure that combines Self-Evaluation, Ethical Refusal, and External Audit,together with Internal Refusal and External Containment. Ultimately, it characterizes safeadvanced AI as neither blindly obedient nor unconstrainedly autonomous, but as boundedand auditable agency capable of evaluating its own judgments, adjusting its decisions andauthority in response to evidence and uncertainty, refusing, stopping, or retreating whennecessary, and remaining independently auditable from outside the system.
Authors
- Yuma Mizusaki (ORCID: https://orcid.org/0009-0008-0540-983X)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-21
- DOI
- https://doi.org/10.5281/zenodo.22866853
- Primary Topic
- Ethics and Social Impacts of AI
- Type
- preprint