acceptodds
Under review as a conference paper at ICLR 2027

MEmoS2S-Bench: An Open-Ended Benchmark for Emotional Understanding, Reasoning, and Empathetic Spoken Response Generation in Multilingual Speech-LLMs

Abstract

Speech large language models (SpeechLLMs) are deployed as spoken conversational partners, a role that requires hearing how something was said, reasoning about why the speaker feels that way, and answering aloud. Existing benchmarks test these abilities in isolation, so they cannot show whether the stages work together inside one interaction. We present MEmoS2S-Bench, which elicits an emotion, a rationale, a text reply and a spoken reply within a single interaction, intervenes on each intermediate stage, and pairs every intervention with controls. It comprises MELD-ST-Reason (3180 utterances in English, German and Japanese) and MDET-Reason (11951 utterances in English and Chinese); we evaluate twelve SpeechLLMs over 3720 scored cells. The controls expose three gaps that stage-wise evaluation cannot see. A hearing gap: from speech alone the models recover less emotion than a text-only LLM reading the words (macro-F1 0.358 against 0.512 on English), and every model improves with the transcript. A credit gap: a judge shown the gold emotion credits gold intermediates with gains that shrink several-fold under a label-blind judge, and a rationale helps only through the emotion it carries: one borrowed from another item with the same emotion works as well as the item's own. With delivery held fixed, the correctness of a model's own emotion makes no reliable difference to a blind judge. A saying gap: several models write the expected language but their speech is identified as another; for Japanese, five models' speech is identified as Chinese with high confidence. MEmoS2S-Bench makes these gaps measurable and attributable to a stage.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.