acceptodds
Under review as a conference paper at ICLR 2027

ARGeoSplat: Autoregressive Geometry Generation and Refinement for Sparse-View 3D Reconstruction

Abstract

A few posed photographs leave much of a scene unseen. Multiview generation can fill in missing views, but plausible RGB images alone do not specify a renderable 3D scene. We extend a next-scale autoregressive multiview generator to produce RGB images and camera-frame pointmaps together. Separate bitwise tokenizers and prediction heads represent appearance and geometry, while a shared transformer conditions each scale on both modalities at earlier scales. We train the geometry predictions with decoded XYZ supervision through a straight-through estimator and a ground-truth-guided cross-view consistency loss. The generated pointmaps then provide proposals for a feed-forward Gaussian predictor. A multiview cost volume searches near each proposed depth to correct local errors, and training on free-running generator outputs exposes the predictor to the errors it will see at inference. With two or three input views, our Gaussian renders exceed EscherNet in PSNR by 1.47–4.02 dB on GSO and OO3D and reduce cross-view photometric error by 0.16–0.24. In ablations, geometry supervision lowers XYZ L2 from 0.0194 to 0.0144, while local depth refinement raises render PSNR from 21.11 to 23.12.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.