acceptodds
Under review as a conference paper at ICLR 2027

ScenePlacer: Reformulating Scene-Level Object Pose Estimation as Visibility Field Learning

Abstract

We present ScenePlacer, a framework for scene-level pose estimation of objects generated by state-of-the-art native shape generation models, placing generated 3D objects into scenes given a partial geometric observation without modifying the underlying generative pipeline. Existing scene generation approaches either generate objects directly in scene coordinates, compromising their canonical-space priors, or regress poses from complete generated shapes in the shape generation process — the latter being a poorly conditioned problem that is sensitive to scale ambiguity and partial observations. ScenePlacer decouples object generation from pose estimation by reformulating alignment as conditional visibility-field learning. Given an observed partial point cloud and a generated mesh, it predicts a visibility field defined in object space, then estimates the rigid transformation by aligning the partial observation to visible parts on the mesh. We propose a pretrained 3D shape VAE as our backbone, preserving its mesh prior while conditioning it on a separately encoded observation. We further introduce visibility-aware, radius-constrained supervision and synthesize training pairs using only object-level meshes and rendered views, eliminating the need for scene-level pose annotations. On 3D-FUTURE, ScenePlacer consistently outperforms existing placement and pose-refinement baselines under both ground-truth and monocular depth estimation. With generated objects, it achieves a Chamfer distance of 0.0077 and an F-score of 92.40 under estimated depth and improves to 0.0023 and 98.27, respectively, under ground-truth geometry.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.