Decodable, Compositional Affect-Like Representations in a Self-Taught Agent and Pretrained Language Models

We study whether emotion-like internal variables arise without hand-coding in artificial networks, and whether affect-related representations share common geometric structure across systems. In a self-taught homeostatic reinforcement-learning agent - trained only with a well-being reward and no emotion labels or objectives - affect-related information forms a decodable, compositional, low-dimensional subspace. When interoceptive variables are hidden from the input, externally grounded affect-like states are still internally constructed. Unsupervised analysis recovers a dominant valence axis together with distinct threat and novelty axes. This geometry is largely generic to structured latent states; what survival makes specific is the valuation of those states. Hidden-level causal control is weak but specific, while interventions at the interoceptive level produce stronger behavioral effects. In pretrained language models, affect directions show alignment with human valence judgments across 1,500 held-out Warriner words (Pearson r = 0.65, 95% CI [0.62, 0.68]) and can be causally steered along valence and arousal dimensions. Phenomenal experience is neither claimed nor measurable. Negative results - including a refuted viability-gradient hypothesis and weak hidden-level causal control - are reported explicitly. Turkish and English versions are included.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-16
DOI
https://doi.org/10.5281/zenodo.22791990
Primary Topic
Emotion and Mood Recognition
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

Decodable, Compositional Affect-Like Representations in a Self-Taught Agent and Pretrained Language Models

Yusuf Asan
Zenodo (CERN European Organization for Nuclear Research)
Emotion and Mood Recognition
preprint

Decodable, Compositional Affect-Like Representations in a Self-Taught Agent and Pretrained Language Models

Yusuf Asan
preprint en

Abstract

We study whether emotion-like internal variables arise without hand-coding in artificial networks, and whether affect-related representations share common geometric structure across systems. In a self-taught homeostatic reinforcement-learning agent - trained only with a well-being reward and no emotion labels or objectives - affect-related information forms a decodable, compositional, low-dimensional subspace. When interoceptive variables are hidden from the input, externally grounded affect-like states are still internally constructed. Unsupervised analysis recovers a dominant valence axis together with distinct threat and novelty axes. This geometry is largely generic to structured latent states; what survival makes specific is the valuation of those states. Hidden-level causal control is weak but specific, while interventions at the interoceptive level produce stronger behavioral effects. In pretrained language models, affect directions show alignment with human valence judgments across 1,500 held-out Warriner words (Pearson r = 0.65, 95% CI [0.62, 0.68]) and can be causally steered along valence and arousal dimensions. Phenomenal experience is neither claimed nor measurable. Negative results - including a refuted viability-gradient hypothesis and weak hidden-level causal control - are reported explicitly. Turkish and English versions are included.

Zenodo (CERN European Organization for Nuclear Research)
Istanbul University-Cerrahpaşa (TR)
Emotion and Mood Recognition
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.