acceptodds
Under review as a conference paper at ICLR 2027

PianoBench: Evaluating Musical Perception and Reasoning in Language Models

Abstract

Multimodal language models continue to improve on broad music understanding benchmarks, but which musical features can they perceive and reason about? We investigate this question by introducing PianoBench, a benchmark that tasks language models with reasoning about musical performances across congruent acoustic and symbolic input representations. PianoBench is derived from our new multimodal, training-scale dataset comprising more than 20,000 hours of interleaved transcripts of music and spoken musical commentary from instructional piano videos. On PianoBench, none of the open-weight audio-language models we evaluate significantly outperform a no-audio control, while proprietary frontier models show mixed gains from audio. In contrast, frontier models using chain of thought achieve up to 93.2% accuracy when provided transcripts of music as text, exposing a gap between reasoning about musical events in symbolic form and perceiving them from audio. To investigate this further, we mid-train Qwen3.5-9B-Base on our interleaved dataset to improve use of music transcripts without chain of thought, raising accuracy on PianoBench from 42.8% to 56% after instruction tuning. These experiments suggest that instructional videos may represent an overlooked source of multimodal supervision for music. Supplement: https://anon834910.github.io/anon-dataset-demo/.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.