Modality-decoupled Distribution Matching Distillation for Few-step Audio-visual Generation
Abstract
Distribution matching distillation (DMD) is an effective route to few-step generation, but its standard formulation treats the audio and video outputs of a joint generator as if they shared the same noise geometry. We find that this coupling is harmful in audio-visual distillation: the two modalities prefer different renoising scales, while a joint fake-score objective can underfit one modality even when its aggregate loss appears well behaved. We introduce , a modality-decoupled DMD formulation with two changes. First, it uses separate audio and video noise schedules while retaining in-distribution joint model evaluations. Second, it trains the fake score with branch-specific objectives at the noise scale used by each modality. On a four-step LTX-2.3 audio-visual generator, these changes improve the normalized full VBench score from 0.642 to 0.674 and the CLAP score from 0.268 to 0.306. These results identify noise-scale and fake-marginal mismatch as central obstacles in few-step joint audio-visual distillation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.