acceptodds
Under review as a conference paper at ICLR 2027

EmoR1: Neutral-Calibrated Reinforcement Learning for Multimodal Emotion Recognition

Abstract

Multimodal Emotion Recognition (MER) combines linguistic, acoustic, and visual cues. Reinforcement-learning-based multimodal large language models (MLLMs) improve coarse-grained emotion and polarity prediction, yet two reliability issues remain. Neutral samples may be forced into non-neutral labels (neutral under-prediction), while non-neutral samples may collapse into neutral (neutral over-prediction), distorting the boundary in opposite directions. Modality-aware rewards may also credit mentions of audio or visual cues that do not improve prediction. Aggregate evaluations obscure these failures when they do not isolate neutral performance or test modality contributions. We introduce the Multimodal Affective Reasoning-Integrated dataset (MARI), a 103K-sample corpus with step- wise reasoning chains and cross-model auditing, and EmoR1, a 7B model trained on MARI with NH-GRPO. NH-GRPO combines bidirectional neutral-boundary rewards with paired full-modal and text-only score comparisons. We also introduce NH-MERBench for evaluating exact labels, semantic clusters, and polarity under a strict neutral criterion. On seven public datasets, EmoR1 achieves competitive NH-MERBench performance and supports further study of neutral calibration and modality credit.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.