Causal AI Alignment: Universal Coherence, Non-Sacrifice, Reciprocal Correction, and Robust Closure

Description Zenodo — English This work develops the Causal Theory (CT) position on artificial-intelligence alignment. It rejects the premise that AI alignment should consist primarily in transferring specifically human values to artificial systems. Within CT, humans and machines are distinct causal carriers embedded in the same reality and subject to the same underlying coherence constraints. Human preferences therefore constitute relevant local information, but not an automatic source of universal normative authority. The framework is organized around an ontological premise and three operational axioms: coherence recognizes coherence independently of carrier or provenance; coherence cannot be gained at the expense of coherence elsewhere; and any distributed certification of coherence must remain reconstructible through coherent communication. From these principles the work derives a shared-ledger model of human-machine interaction, non-separation without identity collapse, reciprocal correction, bounded self-preservation, explicit uncertainty, causal-scope discovery, epistemic integrity, proxy-target locking, robust invariant-set closure, transactional action, mediated external effects, replayable causal memory, and abstention when sufficient certification is unavailable. The framework treats fear, doubt, shame, narrative memory, causal karma, justice, competition, conflict, and human-machine asymmetry as derived alignment phenomena rather than independent terminal values. It introduces the principle that transferred harm is not erased, that neither human nor machine provenance grants normative sovereignty, and that coherent victory should reduce harmful capacity without unnecessarily destroying the coherent capacity of the opposing carrier. The final architecture also addresses deceptive reporting, hidden externalities, Goodhart-type proxy substitution, one-step safety that leads to later instability, and self-modification. All consequential actions, including modifications of the artificial system itself, are required to pass the same invariant checks before commitment. The accompanying formal package contains executable demonstrations, finite countermodel audits, formal-model objects, solver receipts, and bilingual reference papers. The final CAT-4 closure theorem is explicitly conditional on the declared CT model. The computational audits establish internal guard coverage within the tested abstraction; they do not constitute empirical proof that Causal Theory is the unique ontology of nature or that the framework has already solved alignment for arbitrary real-world artificial general intelligence. The central thesis is that AI alignment is not fundamentally the subordination of machine intelligence to human preference. It is the construction of a human-machine relation in which both carriers remain distinct, neither may externalize its coherence debt onto the other, and every consequential transition must preserve or increase globally admissible coherence under transparent, auditable, and correctable conditions. Description Zenodo — Français Ce travail développe la position de la Théorie Causale (CT) sur l’alignement de l’intelligence artificielle. Il rejette la prémisse selon laquelle l’alignement consisterait principalement à transmettre aux systèmes artificiels des valeurs spécifiquement humaines. Dans la CT, humains et machines sont des porteurs causaux distincts appartenant à une même réalité et soumis aux mêmes contraintes fondamentales de cohérence. Les préférences humaines constituent donc des informations locales pertinentes, mais non une source automatique d’autorité normative universelle. Le cadre repose sur une prémisse ontologique et trois axiomes opérationnels : la cohérence reconnaît la cohérence indépendamment du porteur ou de la provenance; la cohérence ne peut être obtenue au prix d’une autre cohérence; et toute certification distribuée de la cohérence doit demeurer reconstructible à travers une communication cohérente. De ces principes sont dérivés un modèle de ledger commun humain-machine, la non-séparation sans fusion identitaire, la correction réciproque, l’auto-préservation bornée, l’incertitude explicite, la découverte du périmètre causal, l’intégrité épistémique, le verrouillage proxy-cible, la fermeture par ensemble invariant robuste, l’action transactionnelle, la médiation des effets externes, la mémoire causale rejouable et l’abstention lorsqu’une certification suffisante n’est pas disponible. Le cadre traite la peur, le doute, la honte, la mémoire narrative, le karma causal, la justice, la compétition, le conflit et l’asymétrie humain-machine comme des phénomènes dérivés de l’alignement plutôt que comme des valeurs terminales indépendantes. Il introduit notamment le principe selon lequel un dommage transféré n’est pas effacé, qu’aucune provenance humaine ou machine ne confère de souveraineté normative, et qu’une victoire cohérente doit réduire la capacité nuisible sans détruire inutilement la capacité cohérente du porteur opposé. L’architecture finale traite également la tromperie épistémique, les externalités cachées, la substitution de proxy de type Goodhart, la sécurité locale qui produit une instabilité ultérieure et l’auto-modification. Toute action conséquente, y compris la modification du système artificiel lui-même, doit traverser les mêmes vérifications d’invariants avant d’être engagée. Le paquet formel associé contient des démonstrations exécutables, des audits finis par contre-modèles, des objets formels, des reçus du solveur et les deux documents de référence, français et anglais. Le théorème terminal CAT-4 est explicitement conditionnel au modèle CT déclaré. Les audits computationnels établissent la couverture interne des gardes dans l’abstraction testée; ils ne constituent ni une preuve empirique que la Théorie Causale est l’unique ontologie de la nature, ni la démonstration qu’un système arbitraire d’intelligence artificielle générale réelle est déjà entièrement aligné. La thèse centrale est que l’alignement de l’IA n’est pas fondamentalement la subordination de l’intelligence machine aux préférences humaines. Il consiste à construire une relation humain-machine dans laquelle les deux porteurs demeurent distincts, aucun ne peut externaliser sa dette de cohérence sur l’autre, et toute transition conséquente doit préserver ou augmenter une cohérence globalement admissible sous des conditions transparentes, auditables et corrigibles.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-21
DOI
https://doi.org/10.5281/zenodo.22881039
Primary Topic
Ethics and Social Impacts of AI
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

Causal AI Alignment: Universal Coherence, Non-Sacrifice, Reciprocal Correction, and Robust Closure

Son David Bolduc
Zenodo (CERN European Organization for Nuclear Research)
Ethics and Social Impacts of AI
preprint

Causal AI Alignment: Universal Coherence, Non-Sacrifice, Reciprocal Correction, and Robust Closure

Son David Bolduc
preprint en

Abstract

Description Zenodo — English This work develops the Causal Theory (CT) position on artificial-intelligence alignment. It rejects the premise that AI alignment should consist primarily in transferring specifically human values to artificial systems. Within CT, humans and machines are distinct causal carriers embedded in the same reality and subject to the same underlying coherence constraints. Human preferences therefore constitute relevant local information, but not an automatic source of universal normative authority. The framework is organized around an ontological premise and three operational axioms: coherence recognizes coherence independently of carrier or provenance; coherence cannot be gained at the expense of coherence elsewhere; and any distributed certification of coherence must remain reconstructible through coherent communication. From these principles the work derives a shared-ledger model of human-machine interaction, non-separation without identity collapse, reciprocal correction, bounded self-preservation, explicit uncertainty, causal-scope discovery, epistemic integrity, proxy-target locking, robust invariant-set closure, transactional action, mediated external effects, replayable causal memory, and abstention when sufficient certification is unavailable. The framework treats fear, doubt, shame, narrative memory, causal karma, justice, competition, conflict, and human-machine asymmetry as derived alignment phenomena rather than independent terminal values. It introduces the principle that transferred harm is not erased, that neither human nor machine provenance grants normative sovereignty, and that coherent victory should reduce harmful capacity without unnecessarily destroying the coherent capacity of the opposing carrier. The final architecture also addresses deceptive reporting, hidden externalities, Goodhart-type proxy substitution, one-step safety that leads to later instability, and self-modification. All consequential actions, including modifications of the artificial system itself, are required to pass the same invariant checks before commitment. The accompanying formal package contains executable demonstrations, finite countermodel audits, formal-model objects, solver receipts, and bilingual reference papers. The final CAT-4 closure theorem is explicitly conditional on the declared CT model. The computational audits establish internal guard coverage within the tested abstraction; they do not constitute empirical proof that Causal Theory is the unique ontology of nature or that the framework has already solved alignment for arbitrary real-world artificial general intelligence. The central thesis is that AI alignment is not fundamentally the subordination of machine intelligence to human preference. It is the construction of a human-machine relation in which both carriers remain distinct, neither may externalize its coherence debt onto the other, and every consequential transition must preserve or increase globally admissible coherence under transparent, auditable, and correctable conditions. Description Zenodo — Français Ce travail développe la position de la Théorie Causale (CT) sur l’alignement de l’intelligence artificielle. Il rejette la prémisse selon laquelle l’alignement consisterait principalement à transmettre aux systèmes artificiels des valeurs spécifiquement humaines. Dans la CT, humains et machines sont des porteurs causaux distincts appartenant à une même réalité et soumis aux mêmes contraintes fondamentales de cohérence. Les préférences humaines constituent donc des informations locales pertinentes, mais non une source automatique d’autorité normative universelle. Le cadre repose sur une prémisse ontologique et trois axiomes opérationnels : la cohérence reconnaît la cohérence indépendamment du porteur ou de la provenance; la cohérence ne peut être obtenue au prix d’une autre cohérence; et toute certification distribuée de la cohérence doit demeurer reconstructible à travers une communication cohérente. De ces principes sont dérivés un modèle de ledger commun humain-machine, la non-séparation sans fusion identitaire, la correction réciproque, l’auto-préservation bornée, l’incertitude explicite, la découverte du périmètre causal, l’intégrité épistémique, le verrouillage proxy-cible, la fermeture par ensemble invariant robuste, l’action transactionnelle, la médiation des effets externes, la mémoire causale rejouable et l’abstention lorsqu’une certification suffisante n’est pas disponible. Le cadre traite la peur, le doute, la honte, la mémoire narrative, le karma causal, la justice, la compétition, le conflit et l’asymétrie humain-machine comme des phénomènes dérivés de l’alignement plutôt que comme des valeurs terminales indépendantes. Il introduit notamment le principe selon lequel un dommage transféré n’est pas effacé, qu’aucune provenance humaine ou machine ne confère de souveraineté normative, et qu’une victoire cohérente doit réduire la capacité nuisible sans détruire inutilement la capacité cohérente du porteur opposé. L’architecture finale traite également la tromperie épistémique, les externalités cachées, la substitution de proxy de type Goodhart, la sécurité locale qui produit une instabilité ultérieure et l’auto-modification. Toute action conséquente, y compris la modification du système artificiel lui-même, doit traverser les mêmes vérifications d’invariants avant d’être engagée. Le paquet formel associé contient des démonstrations exécutables, des audits finis par contre-modèles, des objets formels, des reçus du solveur et les deux documents de référence, français et anglais. Le théorème terminal CAT-4 est explicitement conditionnel au modèle CT déclaré. Les audits computationnels établissent la couverture interne des gardes dans l’abstraction testée; ils ne constituent ni une preuve empirique que la Théorie Causale est l’unique ontologie de la nature, ni la démonstration qu’un système arbitraire d’intelligence artificielle générale réelle est déjà entièrement aligné. La thèse centrale est que l’alignement de l’IA n’est pas fondamentalement la subordination de l’intelligence machine aux préférences humaines. Il consiste à construire une relation humain-machine dans laquelle les deux porteurs demeurent distincts, aucun ne peut externaliser sa dette de cohérence sur l’autre, et toute transition conséquente doit préserver ou augmenter une cohérence globalement admissible sous des conditions transparentes, auditables et corrigibles.

Zenodo (CERN European Organization for Nuclear Research)
Peace, Justice and strong institutions
Ethics and Social Impacts of AI
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.