acceptodds
Under review as a conference paper at ICLR 2027

On-Policy Distillation of Video and Multi-View Priors for Multi-View Video Generation

Abstract

Multi-view video generation aims to synthesize synchronized videos of the same dynamic scene from different viewpoints. Direct generation methods face a shortage of synchronized training data, while combining monocular videos and static multi-view images provides only separate supervision for motion and geometry. Geometry-guided methods rely on estimated depth or 3D proxies, whose errors can compromise cross-view consistency. We propose MultiLive, an on-policy dual-teacher distillation framework that combines separately learned video and multi-view priors in a single generator. Our key idea is to concurrently match the temporal and multi-view distributions of the student's own generations to these complementary priors. The video teacher guides motion within each monocular video, while the multi-view teacher guides geometry and appearance across synchronized views. Both distribution-matching objectives update the student using the same generated sample, encouraging the two capabilities to remain compatible as the scene evolves without explicit 3D reconstruction. Extensive experiments demonstrate that MultiLive achieves high visual quality, strong cross-view consistency, and coherent temporal dynamics without multi-view video training data, using only 4 denoising steps at inference.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.