acceptodds
Under review as a conference paper at ICLR 2027

Context or Face? Locating Emotion Representations in CLIP with Sparse Autoencoders

Abstract

Vision-language models such as CLIP can associate emotions with both facial expressions and contextual cues, but their predictions do not reveal which visual evidence contributes to these judgments. We investigate how joy and sadness are represented across facial and contextual evidence in CLIP using sparse autoencoders (SAEs), causal feature ablation, spatial localization, and cross-carrier linear probing. Using FindingEmo decision-factor annotations and cross-fitted feature discovery, we identify context-associated SAE features for both emotions with distinct evidential profiles. Sadness feature 9960 shows reproducible held-out context selectivity, with 95% bootstrap confidence intervals excluding zero across all three folds. Joy feature 7406 shows consistent contextual enrichment and positive heldout selectivity, although its confidence intervals include zero. Both features rank above all 200 randomly sampled SAE control features in context selectivity. At the dense-representation level, a linear probe trained to distinguish joy from sadness using only face-driven images transfers strongly to context-driven images, achieving 0.851 ± 0.011 balanced accuracy and 0.930 ± 0.006 AUROC. These results show that context-associated behavior at the level of individual sparse features can coexist with emotion-discriminative information that remains linearly accessible across facial and contextual evidence.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.