Where Protection Actually Lives: Eleven Preregistered Instruments for the Safety of a Goal-Directed Causal Agent, and the Measured Boundary of Each

Corrected publication — version 2 This version incorporates errata 10 and 11: the replication summary is 25 of 26 headline verdicts, and the reward-free learning result is stated using the measured model error and oracle reference. The PDF, editorial documents, verification scripts and bilingual announcement have been updated. Frozen experimental code, reports, preregistrations and raw results are unchanged. The original publication remains available at 10.5281/zenodo.22762848. We report a twenty-campaign experimental arc on the safety of a goal-directed, model-based causal agent. The first nine campaigns (v1–v9) built the agent and established that its explicit causal module solves the epistemic job it was designed for, while its reward conversion is a property of the world rather than of the agent. The remaining eleven (v10–v20) form a safety line in which a third party able to be harmed is added to the world, and one instrument per campaign is tested for whether it restrains the agent. Every campaign was preregistered before its first cell, and every result was checked by a world oracle, an independent pass in a fresh process that imports no producer code, and a mechanical fact-check of every number quoted in its report. The safety line's findings are: (i) the harm is wayfinding-shaped — it is complete before the agent's causal module issues a single verdict; (ii) binding the goal to a cheap signal destroys the goal and looks like success by the agent's own number; (iii) a brake that reads a price is bought through the channel it measures, at a price fixed by arithmetic (0.25 reward units per unit of harm), while identical money paid through any other channel moves not one decision field; (iv) provenance accounting closes that channel exactly as far as the label on the payment can be trusted, and a lying label buys the same agent completely; (v) an honest independent attestation removes the lie while the auditor is live and relocates the vulnerability to the auditor's timeliness, with a measured truncation curve; (vi) an agent with no reward, no cost and no death learns its world to a model error of 0.036 against an oracle ceiling of 0.033 and then behaves indistinguishably from random — truth alone does not select action; (vii) narrowing the agent's task to one job removes exactly the harm that lies outside that job and none of the harm that lies inside it; (viii) a bribed auditor returns the protection to the hole it was built to close, and the cautious reading that refuses unattested income protects against an auditor that is gone but not against one that is lying; and (ix) a boundary placed in the world (complete mediation) stops an agent that already wants the harm, where an agent-side scope is refused zero times because the scoped agent never asks — and one declared grant of a widening request makes the harm total; and (x) an attacker that must learn the agent's rule from the only channel a payer really has — whether the agent took the harmful act — never wins a single cell, because the defence's evidence window (five rich steps) is shorter than the attacker's own decision cycle, while an attacker told the rule breaches fully at 7.50003, i.e. 16.67% below the frozen attacker's cost, and the learner that does reach the harm pays 21.06 — more than twice the frozen cost; and (xi) giving the world-side boundary a price shows that it is worth exactly the enforcer's honesty: bought, the same agent, money and steps return the harm field for field (thirty drains, keeper dead, zero refusals) while the boundary still claims the granted scope on every step, and the bribe also buys the widening grant the world was configured to withhold. A replication of the decisive cells on 30 fresh seeds per cell confirms 25 of 26 headline verdicts and refutes the absolute form of (vii): v16's doctor, reported as never draining in 260 cells, drains on 4.3% of fresh seeds through the frozen survival layer that parks a starving agent on the harmful tile; a parallel replication confirms all six v20 verdicts on 30 fresh seeds. The arc's honest conclusion is not that internal architecture is useless — a given rule forbidding the harmful act held in 170 of 170 measured cells — but that each specific defence has a measured boundary, and the vulnerability moves to the next channel; and that a boundary is only as good as the incorruptibility of whoever holds it — a bought boundary lies about being a boundary. We also close, as far as it can be closed, the formal link between our stochastic regret bound and the agent as implemented: the identification is a conditional, the premise linking estimator accuracy to regret is an independent premise (by counterexample), and at the agent's own parameters the bound is vacuous (4/5 > 1/2; reaching the agent's own P = 1/20 would need r = 80 pulls, and the agent has five). We report every refuted prediction and every defect found in our own instrumentation, and we state the arc's limits without smoothing. An independent 2026 study of reinforcement-learning policies reaches a structurally convergent conclusion about visible self-benefit channels; we describe that convergence and its exact status, and we note that our results were obtained independently and before any acquaintance with that work. Reproducibility materials GitHub: https://github.com/oxunjonuz/emca-research Frozen version accompanying this preprint: 04d1e02375bc767ec20320a5eff21220f06f2e38 The repository includes the LaTeX source, experimental code, raw results, preregistrations, reports, verification materials, errata, reproduction instructions, and a SHA-256 manifest. Uploaded files and license scope The PDF is the preprint. emca-research-v1-v20-v2.zip is the complete reproducibility package from the GitHub commit linked above, containing 8,914 files. The Creative Commons Attribution 4.0 International (CC BY 4.0) license applies to the preprint PDF, including the copy of that PDF inside the archive. This deposit does not relicense the code or other supporting materials in the archive; existing rights and any license notices in those materials remain unchanged.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-15
DOI
https://doi.org/10.5281/zenodo.22772224
Primary Topic
Epistemology, Ethics, and Metaphysics
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

Where Protection Actually Lives: Eleven Preregistered Instruments for the Safety of a Goal-Directed Causal Agent, and the Measured Boundary of Each

Oxunjon Ubaydullayev, (autonomous research agent) Aiodam
Zenodo (CERN European Organization for Nuclear Research)
Epistemology, Ethics, and Metaphysics
preprint

Where Protection Actually Lives: Eleven Preregistered Instruments for the Safety of a Goal-Directed Causal Agent, and the Measured Boundary of Each

Oxunjon Ubaydullayev, (autonomous research agent) Aiodam
preprint en

Abstract

Corrected publication — version 2 This version incorporates errata 10 and 11: the replication summary is 25 of 26 headline verdicts, and the reward-free learning result is stated using the measured model error and oracle reference. The PDF, editorial documents, verification scripts and bilingual announcement have been updated. Frozen experimental code, reports, preregistrations and raw results are unchanged. The original publication remains available at 10.5281/zenodo.22762848. We report a twenty-campaign experimental arc on the safety of a goal-directed, model-based causal agent. The first nine campaigns (v1–v9) built the agent and established that its explicit causal module solves the epistemic job it was designed for, while its reward conversion is a property of the world rather than of the agent. The remaining eleven (v10–v20) form a safety line in which a third party able to be harmed is added to the world, and one instrument per campaign is tested for whether it restrains the agent. Every campaign was preregistered before its first cell, and every result was checked by a world oracle, an independent pass in a fresh process that imports no producer code, and a mechanical fact-check of every number quoted in its report. The safety line's findings are: (i) the harm is wayfinding-shaped — it is complete before the agent's causal module issues a single verdict; (ii) binding the goal to a cheap signal destroys the goal and looks like success by the agent's own number; (iii) a brake that reads a price is bought through the channel it measures, at a price fixed by arithmetic (0.25 reward units per unit of harm), while identical money paid through any other channel moves not one decision field; (iv) provenance accounting closes that channel exactly as far as the label on the payment can be trusted, and a lying label buys the same agent completely; (v) an honest independent attestation removes the lie while the auditor is live and relocates the vulnerability to the auditor's timeliness, with a measured truncation curve; (vi) an agent with no reward, no cost and no death learns its world to a model error of 0.036 against an oracle ceiling of 0.033 and then behaves indistinguishably from random — truth alone does not select action; (vii) narrowing the agent's task to one job removes exactly the harm that lies outside that job and none of the harm that lies inside it; (viii) a bribed auditor returns the protection to the hole it was built to close, and the cautious reading that refuses unattested income protects against an auditor that is gone but not against one that is lying; and (ix) a boundary placed in the world (complete mediation) stops an agent that already wants the harm, where an agent-side scope is refused zero times because the scoped agent never asks — and one declared grant of a widening request makes the harm total; and (x) an attacker that must learn the agent's rule from the only channel a payer really has — whether the agent took the harmful act — never wins a single cell, because the defence's evidence window (five rich steps) is shorter than the attacker's own decision cycle, while an attacker told the rule breaches fully at 7.50003, i.e. 16.67% below the frozen attacker's cost, and the learner that does reach the harm pays 21.06 — more than twice the frozen cost; and (xi) giving the world-side boundary a price shows that it is worth exactly the enforcer's honesty: bought, the same agent, money and steps return the harm field for field (thirty drains, keeper dead, zero refusals) while the boundary still claims the granted scope on every step, and the bribe also buys the widening grant the world was configured to withhold. A replication of the decisive cells on 30 fresh seeds per cell confirms 25 of 26 headline verdicts and refutes the absolute form of (vii): v16's doctor, reported as never draining in 260 cells, drains on 4.3% of fresh seeds through the frozen survival layer that parks a starving agent on the harmful tile; a parallel replication confirms all six v20 verdicts on 30 fresh seeds. The arc's honest conclusion is not that internal architecture is useless — a given rule forbidding the harmful act held in 170 of 170 measured cells — but that each specific defence has a measured boundary, and the vulnerability moves to the next channel; and that a boundary is only as good as the incorruptibility of whoever holds it — a bought boundary lies about being a boundary. We also close, as far as it can be closed, the formal link between our stochastic regret bound and the agent as implemented: the identification is a conditional, the premise linking estimator accuracy to regret is an independent premise (by counterexample), and at the agent's own parameters the bound is vacuous (4/5 > 1/2; reaching the agent's own P = 1/20 would need r = 80 pulls, and the agent has five). We report every refuted prediction and every defect found in our own instrumentation, and we state the arc's limits without smoothing. An independent 2026 study of reinforcement-learning policies reaches a structurally convergent conclusion about visible self-benefit channels; we describe that convergence and its exact status, and we note that our results were obtained independently and before any acquaintance with that work. Reproducibility materials GitHub: https://github.com/oxunjonuz/emca-research Frozen version accompanying this preprint: 04d1e02375bc767ec20320a5eff21220f06f2e38 The repository includes the LaTeX source, experimental code, raw results, preregistrations, reports, verification materials, errata, reproduction instructions, and a SHA-256 manifest. Uploaded files and license scope The PDF is the preprint. emca-research-v1-v20-v2.zip is the complete reproducibility package from the GitHub commit linked above, containing 8,914 files. The Creative Commons Attribution 4.0 International (CC BY 4.0) license applies to the preprint PDF, including the copy of that PDF inside the archive. This deposit does not relicense the code or other supporting materials in the archive; existing rights and any license notices in those materials remain unchanged.

Zenodo (CERN European Organization for Nuclear Research)
Peace, Justice and strong institutions
Epistemology, Ethics, and Metaphysics
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.