acceptodds
Under review as a conference paper at ICLR 2027

Generating Visual Sounds The Hard Way

Abstract

We can often infer what an event should sound like simply by watching it. To generate such sounds from video, a model must identify which visual details matter acoustically and use them to guide synthesis. Yet many current video-to-audio models rely on captions for these cues and face a trade-off: removing captions degrades semantic and physical accuracy, while using them can worsen event timing. We introduce TacitV2A, which learns to guide audio generation from visual evidence without captions at inference. Captions serve only as a training signal: caption-guided self-distillation adapts its visual reasoner through a shared frozen audio generator, matching video-conditioned audio-velocity predictions to their stronger caption-conditioned counterparts. This alone lets the reasoner internalize much of the guidance that captions provide. We further refine the audio generator with reinforcement learning to better match paired real recordings. Across FoleyBench, VGGSounder, and FlatSounds, quantitative and human evaluations show that TacitV2A is competitive with caption-conditioned models. On FlatSounds, it outperforms all external baselines in overall physical correctness and approaches its caption-conditioned teacher, matching its event coverage with lower timing error. Interactive results are available in the supplementary material.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.