BEYOND TARGET MATCHING: MEASURING AND MITIGATINGNON-TARGETLEAKAGE INSELECTIVE VIDEO-TO-AUDIOGENERATION
Abstract
Selective video-to-audio (V2A) generation aims to synthesize sounds only for user-specified sources while suppressing other visible sound sources in the scene. Generated audio may contain the requested sound along with sounds from other visible sources. We call this non-target leakage and evaluate source selectivity through both target retention and competing-source suppression. To measure this behavior, we introduce VGG-SourcePair, a paired evaluation benchmark for selective generation in dual-source scenes. Given the same silent video containing two sound sources, we alternately designate each source as the target and jointly evaluate target preservation and non-target leakage, providing a more direct measure of source selection ability. We further propose FocusV2A, which uses the target prompt to select local visual evidence relevant to the requested source before audio generation. A FocusSelector estimates the target relevance of visual features and uses it to modulate spatial attention, yielding a target-focused visual condition. Training proceeds in two stages: the first learns target-aware visual selection, and the second adapts the generator to the resulting visual condition. At inference time, only a silent video and a target prompt are required. Experiments show that FocusV2A achieves stronger source selection performance on VGG-SourcePair while maintaining competitive audio quality and audiovisual synchronization on VGG-MonoAudio.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.