acceptodds
Under review as a conference paper at ICLR 2027

LocalSoundBench: Do Video-to-Audio Models Respond Locally to Visual Edits?

Abstract

Videos with several visible events need each sound to occur with its event. A local edit gives a controlled test: the changed event’s audio should update while sounds elsewhere stay as they were. Global fidelity and synchronization metrics can miss this error because the expected sounds still occur somewhere in the clip. We introduce LocalSoundBench to test whether current video-to-audio (V2A) models can do both. Across 21 released settings, when models produce the requested sound, they often also alter the audio of footage that was never edited. This is not because they cannot produce the affected sounds: when we test these events in isolation and probe the models' internal features, the sound identities are still present. The models instead do not reliably keep each sound at the right time. We therefore propose Regional Flow, which, given the event regions, lets the model see the whole scene but uses information from each event only when generating audio at that event’s time. With Regional Flow, interval accuracy rises by 5.9 to 39.1 points on three continuous-latent generators. Regional Flow also improves both edit accuracy and joint success on paired edits for all three models. For MMAudio, training with sounds aligned to their event times and supplying annotated sound and time pairs on continuous video both give gains in the same direction. These results suggest that when a model generates one soundtrack from the full video, scene context alone is not enough: the model needs explicit control over when each sound is produced. Our code is available in the Supplementary Material.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.