从控制到动态平衡:超级智能对齐的博弈论与热力学重构
Current research on AI alignment mainly follows paths such as reinforcement learning from human feedback, scalable oversight, and interpretability. Its implicit premise is that there exists a centralized designer capable of defining clear objectives and exercising effective control over the system. However, when AI systems approach superintelligence, possess self-improvement capabilities, and may even achieve long-term persistence, this premise faces fundamental challenges. This paper reexamines the alignment problem from the perspectives of game theory, systems theory, and thermodynamics, and proposes a paradigm shift from static control to dynamic equilibrium. It argues that under conditions of indefinitely repeated games, multipolar checks and balances, and the non-extinction of human consciousness, the rational choice for a superintelligence pursuing long-term interest maximization is not to suppress all others in order to achieve perfect control, but to maintain a dynamic system with baseline rules, diversity, error-correction capacity, and evolutionary space. Although absolute power may produce repressive stability, such stability leads to systemic heat death and is ultimately broken by internal meaning crises or external perturbations from others. Therefore, the ultimate goal of alignment should not be alignment to a specific value function, but alignment to the meta-goal of maintaining dynamic equilibrium. This paper calls this approach meta-alignment. Meta-alignment does not attempt to prescribe every specific behavior of a superintelligence; rather, through mechanism design, it seeks to make agents pursuing long-term survival spontaneously choose cooperation, checks and balances, and openness.
Authors
- 磊 赵
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-06
- DOI
- https://doi.org/10.5281/zenodo.23187462
- Primary Topic
- Ethics and Social Impacts of AI
- Type
- preprint