Decodable, Compositional Affect-Like Representations in a Self-Taught Agent and Pretrained Language Models
We study whether emotion-like internal variables arise without hand-coding in artificial networks, and whether affect-related representations share common geometric structure across systems. In a self-taught homeostatic reinforcement-learning agent - trained only with a well-being reward and no emotion labels or objectives - affect-related information forms a decodable, compositional, low-dimensional subspace. When interoceptive variables are hidden from the input, externally grounded affect-like states are still internally constructed. Unsupervised analysis recovers a dominant valence axis together with distinct threat and novelty axes. This geometry is largely generic to structured latent states; what survival makes specific is the valuation of those states. Hidden-level causal control is weak but specific, while interventions at the interoceptive level produce stronger behavioral effects. In pretrained language models, affect directions show alignment with human valence judgments across 1,500 held-out Warriner words (Pearson r = 0.65, 95% CI [0.62, 0.68]) and can be causally steered along valence and arousal dimensions. Phenomenal experience is neither claimed nor measurable. Negative results - including a refuted viability-gradient hypothesis and weak hidden-level causal control - are reported explicitly. Turkish and English versions are included.
Authors
- Yusuf Asan
Institutions
- Istanbul University-Cerrahpaşa (TR)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-16
- DOI
- https://doi.org/10.5281/zenodo.22791990
- Primary Topic
- Emotion and Mood Recognition
- Type
- preprint