Look Where You Splat: Splat Guided Feed-Forward 3D Gaussian Prediction
Abstract
Feed-forward 3D Gaussian Splatting has recently moved from predicting a Gaussian per pixel to predicting Gaussians from a fixed set of learnable tokens, so the number of Gaussians no longer scales with the input images. These token-based feed-forward models reconstruct compact scenes with as few as 2K Gaussians, yet we find that their tokens do not always attend to where their own Gaussians project. We observe that for 31.2% of the tokens, the cross-attention weights do not align with their splat regions, where their predicted Gaussians project in the input images. A misaligned token aggregates image patch features from outside its splat region, leaves part of the region unattended, or both. We hypothesize that a token predicts its Gaussians more accurately when it attends to image patches within its own splat region. To exploit this insight, we propose Gaussian Splat Attention Bias (GSAB), which aligns each token's attention weights to its own splat region in the input images. Because GSAB only adds an attention bias and an intermediate decoding of existing tokens, it applies to any token-based model, which we verify on two that differ in pose assumption and scale: a pose-free model predicting 2K Gaussians and a posed model predicting 262K. GSAB improves PSNR by up to 0.80 dB on the posed model and by 0.71 dB on the pose-free model.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.