acceptodds
Under review as a conference paper at ICLR 2027

SplatFoundry: Voxel-Adaptive Gaussian Decoding with Learned Adam Refinement

Abstract

Feed-forward 3D Gaussian reconstruction replaces costly per-scene fitting with a single network evaluation. However, existing primitive allocations impose an unfavorable choice: pixel-aligned predictors replicate Gaussians across views, while a fixed latent codebook cannot grow with scene extent and complexity. We introduce SplatFoundry, a two-stage framework that combines adaptive 3D allocation with learned Adam refinement. First, we aggregate points and features from a frozen geometric foundation model into depth-adaptive voxel tokens. A Transformer decoder applies self-attention to the concatenation of image-patch and voxel tokens, then decodes a constant number of Gaussians from each occupied voxel token. The resulting primitive count adapts to the occupied 3D space rather than scaling with the number of patch tokens or being capped by a global budget. Second, a state estimator predicts per-Gaussian, per-attribute Adam moments and learning rates from voxel tokens. The differentiable post-optimization fits the Gaussians to the context views, while the refined Gaussians are supervised end-to-end by the standard novel-view objective. Unlike cold-start Adam, this learned initialization produces substantial improvement after only a few refinement epochs. To reduce training memory, we introduce the Accumulated Adam Jacobian (AAJ), whose backward-pass memory is independent of the number of Adam refinement iterations. Across DL3DV and three out-of-distribution datasets, our method improves PSNR by dB and uses fewer Gaussians on average than the strongest prior method in each setting.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.