PixGR: Pixel-Space Generative Decoding for Fast Feed-Forward 3D Reconstruction
Abstract
We present PixGR, an efficient framework that integrates lightweight pixel-space generative modeling directly into feed-forward 3D Gaussian reconstruction. Rather than conditioning generation on an explicit and imperfect reconstruction, PixGR encodes the context observations into a 3D-aware scene latent and generatively reconstructs a global 3D Gaussian representation that jointly supports context and target views. We formulate this process with pixel-space flow matching, enabling generative supervision of 3D reconstruction through differentiable rendering without relying on compressed VAE latents. To retain feed-forward efficiency, PixGR encodes the input scene only once and performs iterative inference with a lightweight DiT decoder. Experiments demonstrate improved reconstruction over regression-based approaches in challenging sparse-view settings, particularly with widely spaced observations and view extrapolation, while maintaining sub-second inference and remaining competitive with substantially heavier pretrained generative models. Code will be publicly available.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.