ZipSplat: Observation-Grounded Scene Tokens for Feed-Forward Gaussian Splatting
Abstract
Feed-forward 3D Gaussian Splatting reconstructs scenes from images in seconds. Pixel-aligned methods keep every Gaussian grounded in an observation, but produce redundant, potentially misaligned Gaussians across views. Token-based methods instead aggregate all views into a fixed set of learned tokens but lack explicit observation support and use a fixed reconstruction budget. We introduce ZipSplat, a method that predicts a compact, observation-grounded scene representation whose size is set at inference. It derives each scene token from an image patch rather than a learned query, and decodes it into freely placed 3D Gaussians. Tokens across views are merged into a compact set whose size is adjustable at inference. However, without geometric supervision, decoded Gaussians can drift from their source patch, and merging can mix tokens from different surfaces. We therefore introduce two geometric losses to ground and shape this representation. A cone loss ties each token's Gaussians to its source patch, and GeoCE aligns token similarity with 3D proximity so that tokens observing the same surface merge together. Building on these observation–geometry associations, we further introduce TokenBA, a token-level bundle adjustment that refines scene tokens and camera poses in a few seconds. ZipSplat requires neither camera poses nor intrinsics, yet outperforms pixel-aligned and token-based baselines on DL3DV and RealEstate10K, even those given ground-truth cameras. It further generalizes zero-shot to Mip-NeRF360 and ScanNet++. Code and trained models will be made public.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.