acceptodds
Under review as a conference paper at ICLR 2027

First-Frame-Guided Video Editing with a Visual Autoregressive Model

Abstract

While text-to-video generative models have advanced video editing, maintaining spatiotemporal consistency in direct video manipulation remains challenging. As a practical alternative to direct video editing, first-frame guided video editing has emerged as a new paradigm, which leverages the superior fidelity of image editing by propagating edited attributes from the first frame to the full sequence. Existing first-frame guided video editing methods are primarily built on diffusion and flow matching models, which struggle to disentangle appearance from motion. To tackle this problem, we propose a novel framework that explicitly exploits the inherent inductive bias of visual autoregressive (VAR) models. Unlike diffusion and flow-matching models that adapt appearance and motion on a shared fixed-resolution latent space, VAR exposes an explicit coarse-to-fine hierarchy, making motion easier to disentangle from the appearance during adaptation. Exploiting this structure, we restrict motion adaptation to low scales while preserving high-scale appearance refinement, achieving high-fidelity first-frame guided video editing with robust source motion preservation. Extensive evaluations demonstrate that our approach achieves leading input-frame consistency while maintaining strong source-motion preservation across both local and global edits.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.