EmoCARE: Empowering Visual Emotion Comprehension via Subjective Adaptation and Canonical Optimization
Abstract
Visual Emotion Comprehension (VEC) aims to infer the emotional responses that visual stimuli may evoke in human observers, which requires reasoning beyond visual content to its affective implications. Although chain-of-thought (CoT) reasoning has shown promise for complex reasoning in multimodal tasks, its application to VEC remains challenging, as existing multimodal large language models (MLLMs) often generate unreliable affective reasoning chains. In this work, we systematically investigate CoT reasoning for VEC and reveal that the major bottleneck lies in the quality of self-generated affective reasoning rather than the CoT paradigm itself. Through fine-grained error analysis, we identify four predominant reasoning failure modes and further highlight the challenge posed by subjective emotion ambiguity. To mitigate this, we introduce VECBench-Subjective, a benchmark with subjective emotion annotations that captures the inherent diversity of human affective judgments. Building on these findings, we propose EmoCARE, which leverages Canonical Affective Rubric rEfinement to enhance the quality of affective reasoning. By transforming fine-grained reasoning errors into instance-specific corrective supervision, EmoCARE enables MLLMs to produce more faithful affective reasoning and achieve improved emotion prediction. Extensive experiments on both VECBench and VECBench-Subjective demonstrate that EmoCARE consistently achieves state-of-the-art performance across multiple MLLMs, validating its effectiveness for both conventional and subjective VEC.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.