acceptodds
Under review as a conference paper at ICLR 2027

Affective Representations of Large Language Models Differ Between AI-Centric and Human-Centric Scenarios

Abstract

Linear probes of large language model hidden states are a useful tool to track model behavior, including states corresponding to emotions. In this work, we take a closer look at the specificity of emotion probes constructed under different stimuli and read-out positions, for seven LLMs. We first investigate whether probes depend on the psychological taxonomy of emotions superimposed onto them, finding that probes trained with disparate frameworks align on conceptually related categories. Second, beyond probing emotional states in 3rd-person text scenarios and in human conversation scenarios, we also probe “AI-centric” scenarios, which are stimuli designed to affect AI systems in particular. Surprisingly, while AI-centric emotional responses are less decodable than human-centric ones at the last token of the prompt, their decodability increases at the last token of the generated response, in five of seven models, unlike the other stimuli. We argue that this indicates these stimuli elicit an emotion-like response that is formed during generation rather than before it. In line with this, probes trained on the three stimulus types recover near-orthogonal directions when trained at the last token of the prompt, but start to align when read at the end of the model's turn. For probing studies of emotion, these results suggest that stimulus domain and read-out position must be reported and controlled, while the taxonomy can be chosen on practical grounds.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.