acceptodds
Under review as a conference paper at ICLR 2027

From Seeing to Hearing: Unified Audio-Video Generation with Progressive Adaptation and Modality Forcing

Abstract

Generating temporally synchronized and semantically coherent audio and video within a unified model remains challenging, particularly when supporting both joint and directional generation. In this work, we propose a multi-stage adaptation framework that extends a pretrained video generator (Wan2.1) into a unified audio–video model. Our framework employs modality-specific branches with bidirectional cross-attention and progressively acquires multimodal generation capabilities through three stages. First, audio-only adaptation equips the pretrained backbone with audio generation capability. Second, joint audio–video fine-tuning establishes cross-modal interaction for synchronized audiovisual generation. Finally, we introduce , a task-dependent adaptation strategy that combines clean modality conditioning, directional information routing, and asymmetric optimization to enable video-to-audio and audio-to-video generation while retaining joint-generation capability. Extensive experiments demonstrate competitive joint audio–video generation, strong bidirectional conditional generation, and preservation of the pretrained video generation capability, highlighting the effectiveness of progressive adaptation and Modality Forcing for unified audiovisual generation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.