DiFSplat: Extrapolative Feed-Forward 3D Gaussian Splatting with Diffusion Feature
Abstract
Feed-forward 3D Gaussian Splatting (3DGS) models enable efficient novel view synthesis but struggle with extreme view extrapolation. Their image-based features (IBFT) rely on source observations and become unavailable in unseen regions, leading to geometric holes and blurry artifacts. We propose DiFSplat, an extrapolative feed-forward 3DGS framework that combines deterministic IBFT with generative diffusion features (DiFT) from a pre-trained 3D-aware diffusion model. We project the initial geometry into novel views and initialize uncovered depths through nearest-neighbor dilation. A frozen diffusion U-Net, conditioned on point renderings, provides features for an auxiliary depth predictor that refines these depths by matching across source and novel views in DiFT space. The refined depths are unprojected to instantiate auxiliary voxels as anchors for unseen geometry. We then fuse IBFT and DiFT on the unified voxel set to predict Gaussian parameters. Experiments show improved extrapolation and zero-shot cross-dataset generalization while maintaining high-fidelity interpolation, demonstrating the benefit of using diffusion priors for both geometry completion and feature fusion.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.