acceptodds
Under review as a conference paper at ICLR 2027

Hierarchical Diffusion for Camera-controlled Autoregressive Video Generation

Abstract

We address the task of long-video, camera-controlled generation from a single image. The dominant autoregressive paradigm suffers from two key limitations: (1) error accumulation over time, as distant frames are conditioned on previously generated outputs rather than ground truth, leading to drift; and (2) ambiguity in how to incorporate camera control. Standard solutions—such as self-forcing and conditioning on explicit 6-DOF poses—either incur significant computational overhead or produce inconsistent scene scale in generated videos. We propose a simple and lightweight framework that combines two models, WanAR and FixAR, built on online reconstruction and hierarchical denoising. To enforce consistency with the input camera scale, we draw from feed-forward novel-view synthesis and reformulate video generation as a warping-and-inpainting problem. To mitigate drift, we introduce a hierarchical denoising strategy in which different stages condition on different temporal contexts. Specifically, we first generate frames conditioned on the immediate past, and then refine them using a second denoiser conditioned on a temporally subsampled long-range history to enforce global consistency. This fine-to-coarse scheme enables each generated frame to effectively condition on the full temporal context without incurring a prohibitive cost. Experiments on RealEstate10k and DL3DV demonstrate that WanAR+FixAR outperforms prior methods for camera-controlled long-video generation, producing results that are both temporally stable and consistent with the scale of the input camera trajectories.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.