acceptodds
Under review as a conference paper at ICLR 2027

V2SA-Bench: A Source-Aware Benchmark for Video-to-Spatial Audio Generation

Abstract

Rapid growth in augmented and virtual reality applications has increased the focus on video-to-spatial audio generation, which conveys immersive cues about objects and their positions in space. However, evaluation efforts remain mono-centric and operate at the global-scene level rather than on individual sources. We introduce V2SA-Bench, a source-aware benchmark for video-to-spatial audio models, built on a human-annotated dataset of 1,274 videos across perspective and 360° videos, focusing on stereo and binaural audio and extending to first-order ambisonics. We define five evaluation dimensions, i.e., semantics, temporal, spatial, acoustics, and perceptual quality, with 12 fine-grained, reference-free, human-aligned metrics. Across 13 baselines, we find that models are strongly center-biased, that spatial alignment is the largest bottleneck, and that performance drops from single to multiple sources. We further find that verifying source presence is essential and that higher perceptual quality does not imply better audio-visual alignment. We hope V2SA-Bench and its interpretable, source-aware metrics offer actionable directions for improving spatial audio generation. We will release our code and benchmark to facilitate future research. The demo webpage is included in the supplementary material.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.