acceptodds
Under review as a conference paper at ICLR 2027

PixWorld: Unifying 3D Scene Generation and Reconstruction in Pixel Space

Abstract

3D reconstruction and generation typically follow separate paradigms: pixel-based regression for reconstruction and latent diffusion for generation. Recent methods seek to unify them in a shared latent space. However, autoencoder compression introduces information loss in both tasks, while latent-space diffusion objectives supervise encoded latents rather than the resulting 3D scenes. To address these limitations, we introduce PixWorld, a single model that unifies 3D reconstruction and generation under a pixel-space diffusion paradigm. By supervising diffusion on rendered images rather than encoded latents, PixWorld eliminates the autoencoder bottleneck and directly optimizes the 3D scene. Furthermore, we introduce PixWorld-Streaming through autoregressive diffusion distillation, extending this unified framework beyond a fixed view budget to support online reconstruction and generation over minute-long sequences. Experiments show that PixWorld outperforms latent-space generation methods and matches state-of-the-art reconstruction, while PixWorld-Streaming retains comparable quality under causal streaming inputs, demonstrating the effectiveness of a unified pixel-space approach.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.