Valence Asymmetries In Language Models: Representation, Behaviour, And Self-Report
Abstract
Emotion-related representations in large language models can influence be- haviour, but whether their internal geometry, behavioural effects, and self-report reflect the same information remains unclear. We investigate these relationships across 171 emotions and 15 models, finding a near-planar valence–arousal ge- ometry that emerges with depth and persists through post-training, with valence consistently more decodable than arousal. Valence also shapes safety-related be- haviour: it couples strongly with refusal (ρ = −0.82), while 83% of multi-turn attack trajectories exhibit negative-valence drift. However, internal representation does not guarantee faithful self-report. Despite near-perfect decoding of induced emotions, models exhibit pronounced valence-dependent reporting asymmetries, which become more apparent under competing emotional inductions, where inter- nal representations and reported states frequently diverge. Together, our findings reveal valence as a recurring organising dimension of affective representations while highlighting fundamental distinctions between what models represent, how those representations influence behaviour, and what models report about their in- ternal states.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.