acceptodds
Under review as a conference paper at ICLR 2027

DynaEmpathy: Benchmarking and Training Speech Language Models for Interactive Empathetic Spoken Dialogue

Abstract

Empathy is a core capability for speech language models, enabling them to respond appropriately to both users' words and vocal signals of affect, urgency, uncertainty, and interpersonal need. Despite progress in spoken-empathy evaluation and alignment, dynamic empathic strategies that adapt to evolving semantic and vocal evidence and user feedback remain insufficiently characterized. To evaluate this dynamic multi-turn empathy in speech language models, we introduce DynaEmpathy, an interactive benchmark grounded in two real-world interaction settings: implicit-affect problem solving and emotion disclosure. Its pipeline combines scenario-grounded seed construction, speech-aware response-contingent user simulation, multi-stage quality control, and trajectory-level evaluation. Conditioned on a private user state and the realized history, the simulator makes assistant decisions consequential for subsequent user behavior and the evaluation context. Our evaluation shows that locally helpful or affectively plausible replies often fail to form a coherent, need-sensitive strategy across turns. To address this gap, we train state-conditioned policies from static turn-level states distilled from simulated interactions. Our responsibility-routed, modality-decomposed GRPO assigns stage-specific semantic rewards to text actions and fidelity and delivery rewards to speech actions, without online user-simulator rollouts. Experiments on the in-domain DynaEmpathy benchmark and out-of-domain spoken-empathy benchmarks show that the method improves the model's empathy in both response content and vocal delivery while retaining aggregate speech-to-text performance on general spoken tasks.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.