acceptodds
Under review as a conference paper at ICLR 2027

OPERA: Object Perception Enhances Single-view 3D Reconstruction

Abstract

Single-view 3D reconstruction is a challenging task in computer vision due to information missing from the single input image. Generative model-based approaches can produce plausible 3D objects from a single image, thanks to data-driven priors learnt from rich and large-scale datasets. However, plausible generation does not guarantee fidelity to the geometry and appearance of the particular input object. Inspired by object perception in human vision, we propose OPERA, a framework that guides multi-view diffusion sampling with pretrained perception models through lightweight alignment modules. These modules are trained independently while both the generative and perception models remain frozen, allowing multiple signals to be combined at inference without joint fusion training. We apply OPERA to two state-of-the-art single-view 3D reconstruction baselines on two standard datasets. We also compare our method with recent image-to-3D models. We provide in-depth analyses of the design choices and their effects across datasets and backbones. Our project page is at https://opera-3d.github.io.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.