acceptodds
Under review as a conference paper at ICLR 2027

Emotion Interpretability in Multimodal Models: Disentangling Behavioral, Representational, and Interventional Evidence

Abstract

Reported emotion circuits controlling expression and emotion neurons shared between speech and faces suggest a shared, steerable emotion representation in multimodal models. This claim bundles three kinds of evidence that benchmark accuracy cannot separate: behavioral (which input carries emotion), representational (whether text-learned emotion is recognized in faces), and interventional (whether it steers generation). We test each separately, comparing transcripts with video frames, classifying faces with emotion prototypes built from transcripts, and steering with neuron groups found without emotion labels or a simple mean-difference direction. Transcripts outperform single frames in three vision-language models by 4.82 to 12.69 percentage points. Encouragingly, text-learned emotion carries over to faces: the prototypes beat a shuffled-label baseline in Qwen2.5-VL-7B and, after feature standardization, in Qwen3-VL-32B, and a text-derived joy direction raises the probability of labeling a face as joy by 30.07 and 36.71 points. At the same time, steering depends on the method and registers more with a classifier than a human reader. Neuron groups give no significant gain across 90 conditions in three text models, whereas the mean-difference direction raises classifier-judged emotion in three of five models, lifting surprise by 69.27 points in Qwen3-32B. A blinded rater judged all 54 sampled mean-difference continuations readable and named the target emotion in 2. We study six open-weight models, 1,411 conversational clips, and 240 acted face clips. Claims of a shared, steerable emotion representation are therefore best tested as three claims, with steering effects confirmed by human readers.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.