Knowing about a body is not learning from one: a pilot study of a language-model agent with artificial interoception
A small language model given an artificial body used what it already knew about bodies, but over ten nights of self-training did not learn anything new from its own. We gave Qwen2.5-1.5B-Instruct an agent body in a simple survival world. The body had energy and integrity, which were reported in text (T), injected along the model's own internal direction for the energy number (V), signalled by a novel random direction after damage (N), and imposed as real computational degradation at low energy (D: head ablation, a shorter reasoning budget, a hotter action choice). Each night the model was fine-tuned with LoRA on its own selected experience ("sleep"). In a preregistered 10-day pilot (sleep vs. no-sleep arms, 16 worlds per day), no measure of development passed its rule. Self-report followed the text channel and ignored the activation channel. The novel channel was not labelled as pain. After sleep, decisions became less sensitive to removing V or N than the base model's were. The agent's inner speech collapsed: by days 7–10, 93–98% of its reasoning steps were the single phrase "I'm sorry, but I can't assist with that." This phrase first appeared under low-energy degradation and was then amplified by self-distillation, not by our data filters. A minimal non-text policy (under 5,000 parameters) trained on the same diet learned to avoid dangerous terrain in 4–5 nights (avoid given cue 0.25 → 0.98). The language model never did (≈ 0.09). The minimal policy, however, never learned to eat when energy was low, which the language model did from the start. We also report five methodological pitfalls for studying machine interoception: a naive "energy" probe that actually reads traces of the model's own degraded text, steering damage driven by vector norm rather than direction, short-context calibration that underestimates it, episode-level selection that penalises the protective action 9 : 1, and self-training loops that converge on refusal. This is the second report from Project Mirror; the first is doi:10.5281/zenodo.23086591.
Authors
- Evgenii Borunov (ORCID: https://orcid.org/0009-0003-4398-840X)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-05
- DOI
- https://doi.org/10.5281/zenodo.23165795
- Primary Topic
- Psychiatry, Mental Health, Neuroscience
- Type
- preprint