acceptodds
Under review as a conference paper at ICLR 2027

AnyScene: Towards Highly Controllable Driving Scene Generation at Anywhere and Beyond

Abstract

Recorded driving datasets cannot exhaustively cover the various road layouts, traffic flows, and rare safety-critical scenarios needed for end-to-end autonomous driving. Scaling synthetic data beyond recorded scenes therefore requires a generator that can faithfully realize user-specified layouts, maintain temporally coherent geometry, and render scenes from flexible viewpoints without matching real-world footage. Existing methods struggle to meet these requirements together: end-to-end video generators offer limited geometric control, while occupancy-guided approaches may deviate from BEV layouts and remain constrained by low-frequency signals, reference frames, or fixed camera rigs. We propose AnyScene, a unified occupancy-centric framework that autoregressively generates semantic occupancy sequences from diverse BEV layouts and uses the resulting geometry to synthesize temporally consistent multi-view driving videos. AnyScene follows user-designed layouts and generalizes zero-shot to layouts from unseen datasets and global maps, while supporting reference-free video synthesis and flexible camera configurations. To support the framework, we build nuCraftv2, a 12 Hz dataset based on nuScenes with synchronized BEV layouts and high-quality dense semantic occupancy. Experiments demonstrate zero-shot occupancy generation across these layout sources and state-of-the-art video generation, together with applications to counterfactual observation generation and sparse-view 3D reconstruction.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.