GaussianDiffuser: Diffusion-based 3D Gaussian Pre-training for Autonomous Driving
Abstract
Self-supervised pre-training has become a fundamental component for advancing 3D scene understanding in autonomous driving. While existing approaches commonly rely on rendering-based reconstruction to enforce geometric consistency, their objectives remain largely confined to camera-observable surface signals, which can make the learned representations overly dependent on directly visible rendering cues. In this paper, we propose GaussianDiffuser, an efficient feedforward framework that integrates diffusion-based denoising into 3D Gaussian Splatting (3DGS) pre-training. Diffusion-based denoising provides a complementary pre-training objective for 3D Gaussian scene representations by regularizing the model to recover clean signals from corrupted latents, encouraging the representation to exploit multi-view context and spatial consistency rather than locally clean rendering cues. To realize this objective, GaussianDiffuser introduces a diffusion-aware Gaussian block for a denoising objective during pre-training, learning enriched spatial representations with minimal architectural modifications. Experimental results on the nuScenes dataset demonstrate that GaussianDiffuser improves depth reconstruction and occupancy prediction, reducing Abs Rel by 17% and improving Occ3D mIoU by 2.4 over previous state-of-the-art methods.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.