acceptodds
Under review as a conference paper at ICLR 2027

BEYOND TARGET MATCHING: MEASURING AND MITIGATINGNON-TARGETLEAKAGE INSELECTIVE VIDEO-TO-AUDIOGENERATION

Abstract

Selective video-to-audio (V2A) generation aims to synthesize sounds only for user-specified sources while suppressing other visible sound sources in the scene. Generated audio may contain the requested sound along with sounds from other visible sources. We call this non-target leakage and evaluate source selectivity through both target retention and competing-source suppression. To measure this behavior, we introduce VGG-SourcePair, a paired evaluation benchmark for selective generation in dual-source scenes. Given the same silent video containing two sound sources, we alternately designate each source as the target and jointly evaluate target preservation and non-target leakage, providing a more direct measure of source selection ability. We further propose FocusV2A, which uses the target prompt to select local visual evidence relevant to the requested source before audio generation. A FocusSelector estimates the target relevance of visual features and uses it to modulate spatial attention, yielding a target-focused visual condition. Training proceeds in two stages: the first learns target-aware visual selection, and the second adapts the generator to the resulting visual condition. At inference time, only a silent video and a target prompt are required. Experiments show that FocusV2A achieves stronger source selection performance on VGG-SourcePair while maintaining competitive audio quality and audiovisual synchronization on VGG-MonoAudio.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.