acceptodds
Under review as a conference paper at ICLR 2027

Modality-decoupled Distribution Matching Distillation for Few-step Audio-visual Generation

Abstract

Distribution matching distillation (DMD) is an effective route to few-step generation, but its standard formulation treats the audio and video outputs of a joint generator as if they shared the same noise geometry. We find that this coupling is harmful in audio-visual distillation: the two modalities prefer different renoising scales, while a joint fake-score objective can underfit one modality even when its aggregate loss appears well behaved. We introduce , a modality-decoupled DMD formulation with two changes. First, it uses separate audio and video noise schedules while retaining in-distribution joint model evaluations. Second, it trains the fake score with branch-specific objectives at the noise scale used by each modality. On a four-step LTX-2.3 audio-visual generator, these changes improve the normalized full VBench score from 0.642 to 0.674 and the CLAP score from 0.268 to 0.306. These results identify noise-scale and fake-marginal mismatch as central obstacles in few-step joint audio-visual distillation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.