acceptodds
Under review as a conference paper at ICLR 2027

Expressing an Emotion Is Not Acting on It: Evaluating Emotion Steering against Random Directions in Static Tasks and Long-Horizon Agents

Abstract

Current evidence that emotion representations steer the behaviour of large language models comes largely from single models, from same-construct outcomes, and from comparisons without random-direction controls. It therefore remains unclear whether steering an emotion vector changes a cross-construct behaviour, whether such a change is specific to the emotion, and whether it persists in long-horizon agents. We present a random-direction-controlled evaluation of emotion steering across four open-weight models, twelve emotions and twelve public cross-construct task categories, extended to long-horizon agent benchmarks. We find that emotion vectors shift option probabilities beyond random directions on many tasks, with model sensitivity differing sharply across models and with signs that do not follow emotion valence; the changes concentrate on answers the model is most confident in. The way the direction is applied matters: additive steering changes the textual emotion readout, whereas gain steering changes answers while rarely expressing the emotion in a neutral continuation probe. In long-horizon agents, steering an emotion direction can shift intermediate actions at individual steps, but these deviations usually do not carry through to a substantial change in the final outcome. These results suggest that emotion steering should be evaluated against random directions and on the behaviour of interest, rather than inferred from emotional expression.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.