MusicListenBench: Can Audio LLMs Hear the Difference Yet?
Abstract
Audio large language models (LLMs) are typically evaluated on song genre, mood, and instrumentation. However, their ability to recognize the same melody across two short clips remains untested. We introduce MusicListenBench, a benchmark comprising paired listening tests for melody, harmony, timbre, and tempo. We generate all audio from symbolic music, ensuring precise ground-truth answers. Each test item comes with two variants. A FLIP variant changes the music, so the answer must change. A STAY variant changes only the sound, such as room echo or hiss, so the answer must stay the same. The two variants have opposite answers, so a model that ignores the audio scores 50%, whatever letter it prefers. Seven of eight open audio LLMs score within 4 points of this floor, and three of them give the same letter to 88% to 99% of items. Of four commercial models, both GPT-Audio models are at chance, and the best, Gemini 2.5 Pro, reaches 74.2%, 12.8 points below human listeners (87.0%). Post-training with GRPO on 10k synthetic clean pairs raises Qwen2.5-Omni from 61.2% to 99.0% on held-out clean pairs, 11.5 points above our listeners, and the gain transfers to a task left out of training. Yet on FLIP/STAY the trained model only reaches the level of our listeners (86.0%). It now misses almost no change (1.9%, against 24.7% for Gemini 2.5 Pro), yet it errs on STAY pairs as often as Gemini does (26.1% against 26.9%). On timbre and rhythm it still changes its answer on 20.8% and 16.8% of pairs where only the sound changed, while our listeners made at most one error on each of these tasks. Nine in ten of these errors come from one change, a quiet background hiss. Training taught the model to ignore EQ changes it never saw (1.4% errors), but not hiss or transposition. Models learn to hear a difference much faster than they learn which differences matter. We release (anonymous) data, code [https://anonymous.4open.science/r/MusicListenBench-F38B/], and an open leaderboard [https://musiclistenbench-review.pages.dev].
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.