acceptodds
Under review as a conference paper at ICLR 2027

Edit in Sync: Unified End-to-End Audio-Visual Editing for Human-Centric Videos

Abstract

Human-centric audio-visual editing requires precise alignment of lip movements, gestures, and expressions with speech. Existing methods process inputs through independent video and audio pipelines, preventing cross-modal interaction between visual editing and audio dubbing. Many of these pipelines further rely on mask-based inpainting. As a result, they suffer from poor lip-audio synchronization, cumulative pipeline errors, and an inherent mask-size dilemma. We present AVE, the first unified end-to-end framework that jointly performs reference-guided person swapping, insertion, and deletion across both audio and video modalities without requiring spatial masks. To overcome data scarcity, we design a synthesis pipeline and introduce AVE-Dataset, the first audio-visual editing dataset to provide 100K pairs with appearance and voice references. By combining Synthetic-to-Real reconstruction with preliminary audio warm-up, our framework transcends synthetic pipeline limits and mitigates audio-video token imbalance. AVE outperforms modality-separate audio-visual editing pipelines in editing realism, temporal consistency, and audio-visual alignment.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.