3MSpeechBench: Measuring Behavioral Consistency of Multimodal LLMs across Speakers and Languages
Abstract
Existing benchmarks for multimodal large language models (MLLMs) with speech input primarily evaluate aggregate task performance, providing limited insight into whether model behavior remains stable when the same content is spoken by different speakers. We present 3MSpeechBench, a multi-speaker, multilingual, multi-task benchmark that holds linguistic content fixed while varying the speaker. Built from four established corpora, it covers spoken multiple-choice QA, spoken visual QA, article summarization, and dialogue summarization, with 50 speakers per language and up to 10 languages. We evaluate MLLMs using both task performance and within-instance consistency across speakers. Our results show that strong average performance does not necessarily imply stable behavior across speakers, highlighting the need to evaluate speech-facing MLLMs beyond aggregate task scores. A demo with data examples is available at https://anonymous.4open.science/w/3mspeechbench-demo-27A3/.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.