ReView: Novel-View Video Generation with Spatiotemporal Consistency
Abstract
We study novel-view video generation: given an input video and a target camera viewpoint, the goal is to synthesize a high-quality video that faithfully replays the same dynamic event from the target view while preserving spatiotemporal consistency with the source video. We introduce ReView, a two-stage framework built on a pretrained video generation model. First, we adapt the model to novel-view generation through supervised training on a mixture of synthesized monocular-video pairs, simulated multi-view videos, and real-world multi-view images. We then apply reinforcement learning (RL) as a post-training stage to further improve spatiotemporal consistency. Specifically, we introduce a Cross-View Reprojection Reward that promotes cross-view geometric consistency through 3D reprojection and visual matching. To improve the efficiency of video RL, we further propose Progressive-Horizon Reinforcement Learning, which progressively shifts updates from single-frame to joint-video rollouts through single-frame, mixed, and full-video stages. This strategy attains a geometric consistency performance comparable to full-video RL using less than one third as many training compute. Experiments show improved viewpoint accuracy and cross-view consistency, with competitive novel-view reconstruction and video quality. These results suggest that pretrained video priors, together with multi-view supervision and geometry-aware post-training, offer a practical route toward more scalable 4D generation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.