X-Splatter: Feed-Forward Metric 4D Gaussian Splatting for Surround-View Driving Scenes
Abstract
Feed-forward scene reconstruction supports scalable simulation and world modeling for autonomous driving. We present X-Splatter, a surround-view, feed-forward 4D Gaussian Splatting (4DGS) framework that processes all surround cameras jointly on a VGGT-Ω backbone, with cross-camera attention throughout the network. It introduces two key components: Depth Anchoring Fine-tuning (DAF), which shifts geometry from relative to metric space, and Static–Dynamic Decomposition Learning (SDDL), which combines dynamic gating, static-promoting regularization, and intermediate-frame supervision to alleviate static–dynamic conflation and enhance motion learning. Following LiDAR-supervised depth adaptation, X-Splatter uses self-supervised motion learning, without bounding boxes, tracklets, or optical-flow labels. Trained and evaluated on all surrounding views, including the side views that share the least content with the front camera, X-Splatter achieves the best full-image and dynamic-region novel-view synthesis and the lowest metric-depth RMSE among the evaluated baselines on the Waymo Open Dataset. It also achieves the lowest cross-view 3D error, indicating better geometric agreement across adjacent cameras. The same model generalizes zero-shot to nuScenes and Argoverse 2, outperforming the evaluated baselines in rendering and metric-depth accuracy on both.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.