Horses for Courses: Breaking Modality Symmetry in Few-Step Audio-Video Diffusion Distillation
Abstract
Unified audio-video diffusion transformers generate synchronized video and audio, but the prohibitive cost of iterative sampling necessitates aggressive few-step distillation. Prevailing distillation practice inherits a modality-symmetric recipe, indiscriminately applying identical distribution matching (DM) objectives and synchronous step schedules across both modalities. In this work, we reveal that this symmetric recipe systematically undermines the relational temporal structure of audio. Through a Fisher information analysis of Gaussian data along the diffusion path, we prove that re-noising suppresses information about cross-time correlations at a quadratic rate in the signal fraction , whereas information about first-order statistics decays only linearly, . Consequently, weak long-range acoustic dependencies receive an attenuated distributional gradient that is readily overwhelmed by the bias of the learned fake score. Teacher re-noising diagnostics confirm that audio loses disproportionately more structure than video at the same signal-to-noise ratio. To address this modality asymmetry, we propose Modality-Aware Asynchronous Distillation (MAAD), which breaks distillation homogeneity by decoupling optimization objectives and step schedules within a single shared student. Exploiting the substantial token-count disparity between modalities, asynchronous denoising inserts audio-only sub-steps that reuse the visual key-value cache, affording dense acoustic refinement at merely 6% additional computation. Student-context trajectory alignment (SCTA) evaluates the teacher's audio velocity directly under the student's detached visual cache, removing the cross-modal conditioning mismatch between the visual context of the target and that of the student. On MiniMax-H3 (33B), MAAD outperforms compared methods on all nine singing metrics, reducing MuQ FAD from 11.37 to 4.58, and matches or exceeds the 50-step teacher on four of six song-quality scores.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.