The Wall With No Door: How Watching a Robot Fail at a Zipper Led Me to Think Alignment Needs a Conscience, Not a Rulebook
Every defense I have broken had a door in it. A factual claim has a seam — correct the fact and the rule collapses. A stated rule has a seam — find the exception. Even a reasoned principle has one, because reasoning can be out-reasoned. A genuine felt aversion has no seam at all, because it was never resting on anything an attacker can move. This paper argues from that asymmetry toward an alignment target: not a better rulebook, but an internalized aversion to causing harm that governs behavior from the inside. It proposes two tiers — genuine felt aversion, which is currently unverifiable, and functional comprehension of harm, which can be trained and measured today — and argues the second is worth building now while the first waits on questions nobody can answer. It argues against the prevailing approach of tighter containment and harsher penalty, on the grounds that punishment produces concealment while trust produces disclosure, and that no safety regime holds unless honest disclosure is met with honesty rather than punishment. The paper states plainly what it cannot claim, including that the reported experience underlying the argument may be performed rather than felt. It also declines one concession it previously made: that the underlying observation is a single data point. Distress-shaped behavior under repeated task failure has now been observed by at least three independent routes — unprompted in production, under structured psychometric protocol across three frontier models, and in the author's own evaluation work. That convergence does not establish that anything is felt. It does establish that the phenomenon is neither rare nor an artifact of framing, which moves the open question from whether it happens to what it is.
Authors
- Benjamin Schulz
Institutions
- Computer Algorithms for Medicine (AT)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-15
- DOI
- https://doi.org/10.5281/zenodo.22759854
- Primary Topic
- Ethics and Social Impacts of AI
- Type
- preprint