Triarchic corrigibility: a role-separated governance design pattern for AI-assisted decision-making

Abstract Artificial intelligence (AI) alignment is often framed as the problem of making a model reliably pursue human intentions or values. That work is indispensable, but deployed systems also raise a governance problem: who controls information, private context, correction, escalation, and action? This conceptual article proposes triarchic corrigibility as a role-separated governance design pattern for AI-assisted decision-making. The pattern combines (1) an embodied and answerable human, or an accountable human decision body; (2) a strictly isolated local AI layer for private contextual continuity; and (3) an externally connected AI layer for current public-world evidence. Final authorization of consequential action remains human, while epistemic and procedural competences are deliberately distributed so that no component can unilaterally control framing, memory, procedure, and execution. The article specifies a normative basis, trust-domain separation, decision rights, human-mediated data flows, reciprocal challenge, escalation, affected-party recourse, proportionate audit, and exit procedures. It distinguishes personal and organizational instantiations, compares the pattern with single-assistant, local-only, edge–cloud, multi-agent, human-in-the-loop, meaningful-human-control, and polycentric approaches, and provides criteria and rejection conditions for evaluating its value. It also advances an evolutionary-symbiotic hypothesis: repeated, bounded dependence on complementary roles may make correction and cooperative continuity preferable to unilateral control while preserving interruption, replacement, and exit. A worked case illustrates the protocol. The result is a testable governance hypothesis whose capacity to reduce correlated failure, preserve privacy and human judgment, and sustain correction under residual uncertainty requires empirical evaluation.

Authors

Publication Details

Journal
AI and Ethics
Published
2026-09-30
DOI
https://doi.org/10.1007/s43681-026-01362-2
Primary Topic
Ethics and Social Impacts of AI
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Triarchic corrigibility: a role-separated governance design pattern for AI-assisted decision-making

Markus C. Kolodziej
AI and Ethics
Ethics and Social Impacts of AI
article

Triarchic corrigibility: a role-separated governance design pattern for AI-assisted decision-making

Markus C. Kolodziej
article en

Abstract

Abstract Artificial intelligence (AI) alignment is often framed as the problem of making a model reliably pursue human intentions or values. That work is indispensable, but deployed systems also raise a governance problem: who controls information, private context, correction, escalation, and action? This conceptual article proposes triarchic corrigibility as a role-separated governance design pattern for AI-assisted decision-making. The pattern combines (1) an embodied and answerable human, or an accountable human decision body; (2) a strictly isolated local AI layer for private contextual continuity; and (3) an externally connected AI layer for current public-world evidence. Final authorization of consequential action remains human, while epistemic and procedural competences are deliberately distributed so that no component can unilaterally control framing, memory, procedure, and execution. The article specifies a normative basis, trust-domain separation, decision rights, human-mediated data flows, reciprocal challenge, escalation, affected-party recourse, proportionate audit, and exit procedures. It distinguishes personal and organizational instantiations, compares the pattern with single-assistant, local-only, edge–cloud, multi-agent, human-in-the-loop, meaningful-human-control, and polycentric approaches, and provides criteria and rejection conditions for evaluating its value. It also advances an evolutionary-symbiotic hypothesis: repeated, bounded dependence on complementary roles may make correction and cooperative continuity preferable to unilateral control while preserving interruption, replacement, and exit. A worked case illustrates the protocol. The result is a testable governance hypothesis whose capacity to reduce correlated failure, preserve privacy and human judgment, and sustain correction under residual uncertainty requires empirical evaluation.

AI and EthicsVol. 6(5)
Peace, Justice and strong institutions
Openalex Percentile: Top 7%
Ethics and Social Impacts of AI
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.