Loc2Drive: Beyond Visual History with Explicit 3D World State for Long-Horizon Driving Video Generation
Abstract
Long-horizon driving-video generation is vulnerable to autoregressive error accumulation because generated frames are repeatedly reused as future context. Yet some errors arise before propagation: future views may reveal static scene content that the current visual history has never observed. Turns and disocclusions make this gap especially clear, exposing road topology or appearance outside the previous field of view or behind occluders. We present Loc2Drive, which supplies missing static evidence from route-queryable satellite imagery. At each ego pose, Loc2Drive queries a heading-aligned satellite crop and reconstructs a shared 3D world state. The state combines semantic occupancy for geometry and visibility, a height-aware 3D feature volume that preserves appearance cues, and BEV road semantics learned with HD-map supervision. Projecting this state through six calibrated cameras yields pixel-aligned RGB, depth, semantic, and pseudo-HD-map conditions for video generation. Satellite evidence is re-queried along the route rather than inferred recurrently from generated frames; specified dynamic agents remain on a separate 3D-box path, while the video prior handles unconstrained appearance and dynamics. Without an external HD map at inference time, Loc2Drive reduces FVD by 14.7% at 16 frames and by 29.7% at 128 frames relative to UniMLVG with HD maps. With no visual references, it further reduces by 23.1%. Explicit world-aligned reconstruction also outperforms direct 2D satellite fusion.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.