acceptodds
Under review as a conference paper at ICLR 2027

Structured Observation Modeling for Feed-Forward 4D Driving Reconstruction

Abstract

Feed-forward 4D driving reconstruction aims to recover coherent dynamic scenes from sparse multi-camera temporal observations, while existing methods still struggle with coherent reconstruction from such sparse observations. Through a data-centric analysis, we find that farther target views receive sparser and more distributed geometric support across cameras and time, while driving datasets exhibit heterogeneous reconstruction conditions, motivating structured observation modeling and heterogeneous training, respectively. In this paper, we propose SOM4D, a feed-forward framework designed to better model sparse multi-view temporal observations for 4D driving scene reconstruction. Specifically, we introduce Pose-Time Relational Adaptation (PTRA) to explicitly model camera–time relations within a frozen visual-geometry backbone. We further introduce Hierarchical Observation Consolidation (HOC) to integrate multi-level features and shared scene context while preserving dense spatial details. To improve explicit scene reconstruction, we develop a Corrective Gaussian Head (CGH) that decodes dynamic Gaussians and progressively corrects them through geometry-aware scene correction and recurrent source-render feedback. Additionally, we train SOM4D on mixed-data to broaden the training conditions and assess the benefit of heterogeneous training on Waymo. Extensive experiments under both Waymo-only and mixed-data settings validate the proposed model design and the benefit of heterogeneous training.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.