EmotionAdapt-R1: Confidence-Calibrated Adaptive Reasoning for Speech Emotion Recognition
Abstract
Speech emotion recognition inherently involves varying ambiguity, ranging from explicit cues to subtle, context-dependent expressions. While recent SpeechLLMs leverage Chain-of-Thought for explainable emotion reasoning, they apply static strategies allocating equal computation regardless of complexity. This fixed-depth method is often inefficient for clear cases and insufficient for ambiguous ones. To address this, we propose EmotionAdapt-R1, a confidence-aware framework enabling adaptive reasoning through Uncertainty-of-Thought, where uncertainty is integrated into generation to realize a "thinking fast and slow" paradigm. Specifically, we introduce Confidence-Adaptive GRPO (CA-GRPO), a policy optimization algorithm that dynamically aligns accuracy, reasoning depth, and confidence expression. Extensive evaluations show our approach outperforms strong baselines on standard SER benchmarks, demonstrating the value of self-aware confidence for efficient, adaptive multimodal reasoning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.