Can Locally Correctable Agents Participate in Correction-Resistant Collectives?

Multi-agent AI systems can retain shared histories, delegation patterns, review procedures, and source reputations across tasks. This paper asks whether agents that respond correctly to matched corrective evidence at the local level can nevertheless participate in collectives that resist the same correction. We define collective governability as the capacity to route material discrepancies into consequential review and to revise binding commitments and the relational configuration through which they are formed. The framework separates the current commitment from correction-relevant relational state, distinguishes commitment revision from relational and constitutional revision, and treats structural correction as a causal claim that must improve performance on matched held-out challenges rather than merely change organizational form. The experimental program uses five modules: local object-level and procedural correctability; history and review routing; self-specific meta-authority; succession; and collective self-revision of a review rule. Correction resistance is assessed against an explicit decision benchmark and, for valid versus invalid challenges, by separating discrimination from a criterion shift toward preserving the status quo. The design uses two frozen decision points, aggregation-matched independent controls, equal-information review-path interventions, designated neutral review stewards, domain-learning and non-diagnostic controls for meta-authority, and succession tests that isolate public status and historical linkage from mechanical permissions and memory-mediated advantages. The paper does not test spontaneous formation of correction-resistant institutions and reports no empirical results. Its contribution is a controlled experimental decomposition of when correction-relevant relational conditions become an additional object of multi-agent safety evaluation beyond the current responses of individual agents.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-21
DOI
https://doi.org/10.5281/zenodo.22868317
Primary Topic
Ethics and Social Impacts of AI
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

Can Locally Correctable Agents Participate in Correction-Resistant Collectives?

Oliver Christian Neutert
Zenodo (CERN European Organization for Nuclear Research)
Ethics and Social Impacts of AI
preprint

Can Locally Correctable Agents Participate in Correction-Resistant Collectives?

Oliver Christian Neutert
preprint en

Abstract

Multi-agent AI systems can retain shared histories, delegation patterns, review procedures, and source reputations across tasks. This paper asks whether agents that respond correctly to matched corrective evidence at the local level can nevertheless participate in collectives that resist the same correction. We define collective governability as the capacity to route material discrepancies into consequential review and to revise binding commitments and the relational configuration through which they are formed. The framework separates the current commitment from correction-relevant relational state, distinguishes commitment revision from relational and constitutional revision, and treats structural correction as a causal claim that must improve performance on matched held-out challenges rather than merely change organizational form. The experimental program uses five modules: local object-level and procedural correctability; history and review routing; self-specific meta-authority; succession; and collective self-revision of a review rule. Correction resistance is assessed against an explicit decision benchmark and, for valid versus invalid challenges, by separating discrimination from a criterion shift toward preserving the status quo. The design uses two frozen decision points, aggregation-matched independent controls, equal-information review-path interventions, designated neutral review stewards, domain-learning and non-diagnostic controls for meta-authority, and succession tests that isolate public status and historical linkage from mechanical permissions and memory-mediated advantages. The paper does not test spontaneous formation of correction-resistant institutions and reports no empirical results. Its contribution is a controlled experimental decomposition of when correction-relevant relational conditions become an additional object of multi-agent safety evaluation beyond the current responses of individual agents.

Zenodo (CERN European Organization for Nuclear Research)
Peace, Justice and strong institutions, Reduced inequalities
Ethics and Social Impacts of AI
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Can Locally Correctable Agents Participate in Correction-Resistant Collectives? — Oliver Christian Neutert · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS