acceptodds
Under review as a conference paper at ICLR 2027

PoG: Point-conditioned Object Generation for Compositional 3D Scene Reconstruction

Abstract

Compositional 3D scene reconstruction seeks to recover complete object-centric assets and their spatial arrangement from images. Native 3D object generation models, which learn powerful shape and appearance priors from large-scale asset collections, have recently enabled the synthesis of complete, textured objects and inspired compositional scene reconstruction from a single image. Yet a single view provides inherently incomplete evidence: occlusion and perspective ambiguity leave both object geometry and spatial arrangement underconstrained. Multi-view images offer complementary observations that can reduce these ambiguities, but leveraging them with canonical object priors requires associating evidence across views and reconciling scene-space observations with the canonical coordinates used by object generators. In this paper, we present PoG, a compositional 3D scene reconstruction method that generates a complete 3D asset for each observed object instance from multi-view RGB images through point-conditioned object generation. Given only multi-view images, an off-the-shelf reconstruction and segmentation frontend establishes cross-view instance observations and derives a partial point cloud for each object; no depth or object geometry is provided as input. PoG augments these points with DINO features and encodes the resulting feature-bearing point set into spatially grounded condition tokens. It then jointly generates canonical object occupancy and predicts the relative similarity transformation between the canonical asset and the normalized observation frame. The predicted transformation canonicalizes the point conditions for subsequent geometry and texture generation and places each completed asset back into the scene, eliminating post-hoc optimization-based registration. Experiments on Toys4K and multi-view 3D-FRONT demonstrate the effectiveness of PoG for object- and scene-level reconstruction.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.