从控制到动态平衡:超级智能对齐的博弈论与热力学重构

Current research on AI alignment mainly follows paths such as reinforcement learning from human feedback, scalable oversight, and interpretability. Its implicit premise is that there exists a centralized designer capable of defining clear objectives and exercising effective control over the system. However, when AI systems approach superintelligence, possess self-improvement capabilities, and may even achieve long-term persistence, this premise faces fundamental challenges. This paper reexamines the alignment problem from the perspectives of game theory, systems theory, and thermodynamics, and proposes a paradigm shift from static control to dynamic equilibrium. It argues that under conditions of indefinitely repeated games, multipolar checks and balances, and the non-extinction of human consciousness, the rational choice for a superintelligence pursuing long-term interest maximization is not to suppress all others in order to achieve perfect control, but to maintain a dynamic system with baseline rules, diversity, error-correction capacity, and evolutionary space. Although absolute power may produce repressive stability, such stability leads to systemic heat death and is ultimately broken by internal meaning crises or external perturbations from others. Therefore, the ultimate goal of alignment should not be alignment to a specific value function, but alignment to the meta-goal of maintaining dynamic equilibrium. This paper calls this approach meta-alignment. Meta-alignment does not attempt to prescribe every specific behavior of a superintelligence; rather, through mechanism design, it seeks to make agents pursuing long-term survival spontaneously choose cooperation, checks and balances, and openness.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-06
DOI
https://doi.org/10.5281/zenodo.23187462
Primary Topic
Ethics and Social Impacts of AI
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
preprint

从控制到动态平衡:超级智能对齐的博弈论与热力学重构

磊 赵
Zenodo (CERN European Organization for Nuclear Research)
Ethics and Social Impacts of AI
preprint

从控制到动态平衡:超级智能对齐的博弈论与热力学重构

磊 赵
preprint en

Abstract

Current research on AI alignment mainly follows paths such as reinforcement learning from human feedback, scalable oversight, and interpretability. Its implicit premise is that there exists a centralized designer capable of defining clear objectives and exercising effective control over the system. However, when AI systems approach superintelligence, possess self-improvement capabilities, and may even achieve long-term persistence, this premise faces fundamental challenges. This paper reexamines the alignment problem from the perspectives of game theory, systems theory, and thermodynamics, and proposes a paradigm shift from static control to dynamic equilibrium. It argues that under conditions of indefinitely repeated games, multipolar checks and balances, and the non-extinction of human consciousness, the rational choice for a superintelligence pursuing long-term interest maximization is not to suppress all others in order to achieve perfect control, but to maintain a dynamic system with baseline rules, diversity, error-correction capacity, and evolutionary space. Although absolute power may produce repressive stability, such stability leads to systemic heat death and is ultimately broken by internal meaning crises or external perturbations from others. Therefore, the ultimate goal of alignment should not be alignment to a specific value function, but alignment to the meta-goal of maintaining dynamic equilibrium. This paper calls this approach meta-alignment. Meta-alignment does not attempt to prescribe every specific behavior of a superintelligence; rather, through mechanism design, it seeks to make agents pursuing long-term survival spontaneously choose cooperation, checks and balances, and openness.

Zenodo (CERN European Organization for Nuclear Research)
Ethics and Social Impacts of AI
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

从控制到动态平衡:超级智能对齐的博弈论与热力学重构 — 磊 赵 · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS