TAP-Foley: Inference-Time Target-Anchored Projection for Counterfactual Video Foley Generation
Abstract
Counterfactual Video Foley Generation replaces the source sound identity that a silent video implies with the target sound identity that a text prompt describes, while preserving the timing of visible events. Video-text-to-audio (VT2A) mod- els struggle with this task because video conditioning carries not only the timing but also the source identity, which conflicts with the target prompt. Existing training-based methods suppress the source identity by adding a visual adapter or by training on video–text conflicts, yet it remains partially audible. Inference-time guidance methods suppress it more fully, by negative source prompting with a fixed weight or by dropping video conditioning at an intermediate step, yet both degrade temporal alignment. We propose TAP-Foley, an inference-time guidance method that adaptively suppresses the source identity while retaining video conditioning throughout generation. At each step, it measures how strongly the video steers the prediction toward the source identity beyond what the target prompt alone would produce. It removes this excess along a single source–target axis with a scaled closed-form projection onto a target-anchored constraint, so the conflict between the two modalities is mitigated without dropping either. To reliably evaluate sound identity replacement and temporal alignment, we introduce VGGSound-Counterfactual, a benchmark for Counterfactual Video Foley Generation. Comprehensive experiments show that TAP-Foley achieves a more favorable trade-off between sound identity replacement and temporal alignment than the evaluated configurations of prior methods. At matched replacement, it improves temporal alignment by 20% over the strongest prior inference-time guidance method. Code, benchmark, and demos are available at https://anonymous129108.github.io/TAP-Foley/
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.