acceptodds
Under review as a conference paper at ICLR 2027

Watching, Not Listening: Benchmarking Omnimodal Vocal Performance Commentary

Abstract

Vocal coaching turns a sung performance into selective, actionable feedback, now within reach of omnimodal large language models (omni models) that watch and listen at once. Omni models have grown into unified audio-visual backbones, while AI for music has moved from understanding to singing assessment. However, vocal commentary needs audio-visual evidence, verifiable expert judgement and selectivity at once, and no benchmark measures all three. We formalize Omnimodal Vocal Performance Commentary (OmniVPC), where a model comments on a sung clip from its video and audio as an expert coach would. We release OmniVPC-Bench, 1,970 items over 127 coached clips in nine genres, with three tracks under four matched media conditions and each coach's own words as the reference. Across 31 omni models, none reaches an OVS of 55: models gain from the picture rather than the sound, and the coach's transcript in place of the media adds about 40 points.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.