EmoInsight: Learning Fine-Grained Visual Affect from Homologous Variations and Preference-Guided Semantic Supervision
Abstract
Visual Emotion Assessment (VEA) aims to model human affective perception elicited by visual content. In real-world scenarios, emotion perception is inherently subjective, arises from reasoning over multiple interacting visual cues, and can vary substantially with subtle changes in visual context. However, existing VEA methods largely treat emotion assessment as numerical scoring. This limits their capacity for deep semantic reasoning over diverse cues and capturing fine-grained visual-affective correlations. To bridge this gap, we introduce a new benchmark dataset (**ContiEmo**) constructed from homologous image groups with densely annotated affective variations and human-annotated multi-cue reasoning labels. Building upon ContiEmo, we develop **EmoInsight**, a preference-guided VEA model. Specifically, a *Multi-Teacher Synergistic Affective Reasoning* mechanism is proposed to enhance semantic reasoning capabilities by transferring diverse affective knowledge with instance-specific dynamic weighting derived from human preferences. To further capture fine-grained visual-affective correlations, we introduce a *Coarse-to-Fine Factor-Guided Reinforcement* strategy that jointly optimizes V-A estimation and consistency with affective factor cues. Extensive experiments on ContiEmo and public real-world VEA benchmarks demonstrate that our method consistently improves fine-grained affect estimation and multi-factor recognition.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.