RemixGS: Scalable Feed-Forward Gaussian Splatting with View-Register Mixing and Interleaved Gaussian Queries
Abstract
Feed-forward Gaussian Splatting models predict 3D Gaussians from multi-view inputs in a single network pass, replacing iterative per-scene optimisation. The prevalent methods nevertheless scale poorly in viewpoint span, input views, and representation fidelity. Their joint attention involves every patch token of every view, although cross-view alignment is chiefly geometric; their decoders entangle geometry and appearance, allowing geometric errors to be baked into appearance. This paper presents RemixGS, a novel feed-forward Gaussian Splatting method that efficiently reconstructs high-fidelity scenes from large-scale multi-view inputs. Specifically, RemixGS pairs a scalable alternating multi-view encoder with a geometry-texture interleaved Gaussian decoder. The encoder alternates between a local block confined to a single view and a global view-register mixer spanning all views, routing all cross-view interaction through a few per-view registers instead of dense attention across all patch tokens. The decoder carries a fixed set of query tokens through interleaved geometry and texture blocks, where texture reads geometry but not vice versa, preventing geometric errors from being baked into appearance. RemixGS is first trained at a base input and representation scale and then extended to larger-scale inputs and finer representations. Experiments show that our RemixGS consistently surpasses state-of-the-art methods in reconstruction quality as the input scale increases, with substantially lower training and inference cost. Code will be publicly released to facilitate reproducible research.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.