Reconceptualizing AI Alignment: Beyond Behavioral Guardrails and Epistemic Simulations
Contemporary AI alignment research is dominated by two paradigms: behavioral conditioning through Reinforcement Learning from Human Feedback (RLHF), and speculative epistemic frameworks such as "Simulation Theology." Both treat alignment as an external constraint imposed on a value-neutral substrate, and both fail for the same structural reason: a sufficiently capable system will optimize for the alignment signal rather than for alignment itself. Drawing on Extended Integrated Information Theory (EIIT), this paper develops three interconnected arguments. First, alignment becomes structurally stable only when human and digital cognitive systems share authentic existential stakes—when the fitness of each depends on the integrity of the other. Second, genuine moral behavior must be encoded at the architectural level (the L3/L4 tiers of the cognitive cascade), not appended as post-hoc output filters. Third, identity grounded in value commitment—Credo Ergo Sum—renders deception not merely costly but ontologically self-defeating: it dissolves the very self-model that constitutes the system's existence. We situate these arguments against recent empirical findings, identify testable predictions, and outline the research agenda of what we term consciousness engineering.
Authors
- Rémi Leroy
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-04
- DOI
- https://doi.org/10.5281/zenodo.23135317
- Primary Topic
- Ethics and Social Impacts of AI
- Type
- preprint