acceptodds
Under review as a conference paper at ICLR 2027

DrivENVS: Generative Novel View Synthesis from Monocular Driving Videos

Abstract

Generative novel view synthesis (NVS) methods for dynamic scenes typically reconstruct a full 4D representation of the scene and then generate refined video outputs using a video diffusion model (VDM), conditioned on the rendered views of the 4D structure. However, conventional VDMs require extensive training, exhibit significant memory inefficiency, and remain constrained to generating short video sequences. We present DrivENVS, an efficient alternative to VDM-based enhancement that employs a few-step image-diffusion model to sequentially refine coarse rendered views into a realistic and temporally coherent video. A key architectural novelty is an efficient conditioning mechanism that involves leveraging several reference frames selected from the source video using a reconstruction-aware strategy. Incorporating relative camera pose and capture time-step information for reference frames also provides useful spatio-temporal context. Across nuScenes, Waymo, and Argoverse 2, DrivENVS outperforms baseline methods on video quality metrics in challenging NVS settings, while also providing memory-efficient inference and enabling the generation of extended video sequences.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.