Intervention, Not Shared Latents: Blocking Visual Shortcuts in Audio-Video Generation
Abstract
Joint audio–video (AV) generators are trained on data in which what an event looks like and what it sounds like are spuriously correlated. We present a controlled causal study of the resulting failure mode. In an AV structural causal model where the audio is, by construction, independent of the video's nuisance appearance, models that let audio read video directly—through cross-attention or a shared latent—learn a visual shortcut: they predict sound from appearance rather than the causal event and, when the appearance–event correlation is broken at test time, synthesize the wrong event's sound. Crucially, the popular remedy of routing both modalities through a shared common-cause latent does not fix this—a bottleneck, an unsupervised shared/private factorization, and a faithful shared-prior model all grab the appearance proxy and fail like the direct model. Blocking the shortcut instead requires an intervention on the nuisance: under the stated assumptions we prove that counterfactual invariance is necessary and sufficient to identify the causal predictor, and we verify the mechanism from feature-vector SCMs to procedural pixel video, real images with spectrogram audio, moving real digits, and a conditional generator. On a real, pretrained V2A generator (MMAudio), an input-intervention test shows the model is far from invariant to sound-irrelevant edits, though a generic-noise control reveals it is broadly input-brittle rather than specifically colour-shortcutting—clean isolation of the shortcut needs the controlled confounds our synthetic studies provide. We characterize when the shortcut occurs, compare the objective against supervised counterfactual augmentation, and isolate the unknown-nuisance regime—where the intervention cannot be applied—as the central open problem.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.