acceptodds
Under review as a conference paper at ICLR 2027

EmoInsight: Learning Fine-Grained Visual Affect from Homologous Variations and Preference-Guided Semantic Supervision

Abstract

Visual Emotion Assessment (VEA) aims to model human affective perception elicited by visual content. In real-world scenarios, emotion perception is inherently subjective, arises from reasoning over multiple interacting visual cues, and can vary substantially with subtle changes in visual context. However, existing VEA methods largely treat emotion assessment as numerical scoring. This limits their capacity for deep semantic reasoning over diverse cues and capturing fine-grained visual-affective correlations. To bridge this gap, we introduce a new benchmark dataset (**ContiEmo**) constructed from homologous image groups with densely annotated affective variations and human-annotated multi-cue reasoning labels. Building upon ContiEmo, we develop **EmoInsight**, a preference-guided VEA model. Specifically, a *Multi-Teacher Synergistic Affective Reasoning* mechanism is proposed to enhance semantic reasoning capabilities by transferring diverse affective knowledge with instance-specific dynamic weighting derived from human preferences. To further capture fine-grained visual-affective correlations, we introduce a *Coarse-to-Fine Factor-Guided Reinforcement* strategy that jointly optimizes V-A estimation and consistency with affective factor cues. Extensive experiments on ContiEmo and public real-world VEA benchmarks demonstrate that our method consistently improves fine-grained affect estimation and multi-factor recognition.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.