From Surface to Depth: Towards Cognitive Appraisal Reasoning in Multimodal Emotion Understanding
Abstract
Recent multimodal large language models (MLLMs) increasingly incorporate explainable reasoning for emotion understanding. However, reasoning based mainly on observable affective cues, without deeper theory-guided analysis of how these cues relate to emotion formation, may reduce emotion understanding to statistical associations between cues and emotion categories. This reliance on superficial associations can give rise to the Clever Hans effectClever Hans Effect: A system exhibits seemingly intelligent behavior but actually relies on hidden cues, data biases, or superficial associations rather than genuinely understanding the mechanisms underlying the task., where model predictions depend on surface-level correlations rather than a genuine understanding of why emotions arise. Consequently, such shortcuts often become unreliable when affective cues are implicit, conflicting across modalities, linguistically misleading, or contaminated by redundant details. %visually absent In contrast, human emotions are shaped by how individuals interpret and evaluate surrounding events beyond observable cues. Inspired by appraisal theories of emotion, we formulate multimodal emotion understanding as a progression from perception to cognitive appraisal, and introduce a dataset, a model, and a benchmark to support this paradigm. 182 CogEmo-40K is a large-scale instruction-tuning dataset constructed through a perception-to-appraisal pipeline, where models generate evidence-grounded reasoning across six appraisal dimensions including goal congruence, coping potential, accountability, future expectancy, arousal, and action tendency. 183 CogEmo-MoE is a sparse MLLM that explores cognitive appraisal-guided emotion reasoning with Mixture-of-Experts architectures. 184 CogEmo-Bench evaluates free-form model outputs by jointly measuring appraisal evidence quality and emotion recognition performance through the Appraisal Evidence Quality Score (AEQS) and emotion classification metrics. Extensive experiments show that our paradigm not only leads CogEmo-Bench (CogEmo-MoE: \(56.30\!\rightarrow\!70.60\) AEQS vs. GPT-5.2; \(63.10%\!\rightarrow\!71.43%\) accuracy and \(52.19%\!\rightarrow\!57.17%\) M-F1 vs. Qwen3-Omni), but also exhibits strong cross-domain generalization, highlighting the effectiveness of perception-to-appraisal reasoning in moving beyond surface-level cue-label associations toward more reliable multimodal emotion understanding.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.