acceptodds
Under review as a conference paper at ICLR 2027

SelFoley: Turning Audio-Conditioned Video-to-Audio Models into Text-Guided Source Selectors

Abstract

Text-guided selective video-to-audio (V2A) asks a model to generate only a requested sound from a silent video that contains several plausible sources. Existing systems learn source selection and soundtrack generation together from large, automatically constructed data. We take a different route. Audio-conditioned V2A (AC-V2A) models normally transfer the content and acoustic character of a reference sound to a new video. In the open-weight ControlFoley model, we find that an isolated reference does more: it suppresses other visible events, and a controlled pathway intervention localizes this choice to an internal vector that represents what sound the reference contains. We call this the source-content condition. The Clean Source Isolation Operator (CSIO) is a 1.31M-parameter adapter that predicts this condition from target text while the generator remains frozen. CSIO alone improves selection but weakens audiovisual timing; SelFoley combines CSIO with early visual-synchronization guidance to recover alignment. We train the adapter on StemFoleyDB-3.3K, a dataset of 3,357 clips and 5,383 human-audited source stems. Its training split contains 3,287 clips (about 9.0 clip-hours); 3,274 clips with complete frozen features supply 300 adapter updates. Relative to SelVA's reported 179k-video, roughly 500-hour task corpus, SelFoley uses 98.2% fewer task-training clips and 96.8% fewer trainable parameters. On a fixed 50-clip test set, SelFoley reaches 78.76% target selection versus 51.33% for SelVA, while its Synchformer error is 0.384 versus 0.413. Eighteen listeners also rate non-target suppression at 3.30 versus 1.98 on a five-point scale. Separately trained CSIO adapters improve target selection by 7.4–23.6 points within three open-weight AC-V2A backbones, showing that generator-specific native reference-content interfaces can be reused instead of relearning separation end to end. Dataset examples and model outputs are available at https://sel-foley.github.io/.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.