DriveReGen: Generating Driving Scenes as Gaussians in Reconstruction Space
Abstract
Collecting real driving data is costly. Video models offer an alternative by generating driving scenes from text and layouts, yet they control these layouts only in 2D space and cannot ensure consistent geometry across frames and views. Recent methods generate Gaussian scenes from generated videos or video latents, but the reconstructed geometry may not fully match the generated appearance. General-scene methods instead generate in feed-forward reconstruction state space to jointly model geometry and appearance. However, conditioning on metric driving layouts requires the frozen decoder's scale and FOVs. We introduce DriveReGen, which generates driving scenes in reconstruction space with scene-specific scale and decoder fields of view for metric layout conditioning. Our two-stage layout control first aligns content with projected HD maps and actor tracks, then uses decoded calibration to encode metric layouts as reconstruction rays and depths. To supervise decoded outputs, Hierarchical Decoding Supervision (HDS) backpropagates feature, geometry, and image errors through frozen decoding. Given text, cameras, and layouts, DriveReGen generates 3D Gaussians without target images or LiDAR. On Waymo, DriveReGen reduces FVD by 58.0% and improves lane F1 by 3.77 percentage points over InfiniCube. Under physical cameras, its road/actor LiDAR AbsRel is 17.0% lower. Code will be made publicly available.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.