acceptodds
Under review as a conference paper at ICLR 2027

FreeView: Multi-View Driving Video Generation with Spatiotemporal Memory

Abstract

Robust autonomous driving requires diverse multi-view observations of common traffic scenarios and long-tail events. Dashcam and online monocular driving videos provide a rich source of these scenarios. Their limited view coverage makes them difficult to use directly in autonomous driving systems that require synchronized multi-camera observations. Expanding these videos into coherent multi-view sequences requires preserving the reference scene across viewpoints and integrating observations from different times. We propose FreeView, a multi-view driving video generation framework with spatiotemporal memory. Our reference-as-view design incorporates the reference video as an additional conditioning view, using shared view-wise self-attention for memory encoding and cross-view attention for memory interaction. This design leverages the strong representation and interaction capabilities of the pretrained generation model. We further retrieve relevant reference observations across time and use relative temporal and spatial encodings to guide attention matching. Experiments demonstrate the effectiveness of our method in preserving visual fidelity and consistency with the reference scene. We also demonstrate generation under flexible camera configurations and applications to real-world dashcam videos.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.