Self-Supervised Object-Centric Cross-View Reasoning from a Single Image
Abstract
We introduce Particle-based Object Representation Transport (PORT), a self-supervised object-centric model that directly transports object-centric latent representations extracted from a single source image to target viewpoints according to relative camera pose. Rather than merely generating novel-view images, PORT directly transports each object latent to predict its position, scale, depth, and appearance under different horizontal camera viewpoints. This enables PORT to construct viewpoint-conditioned multi-view object representations from a single image while preserving object correspondence across views, enabling downstream 3D reasoning by leveraging object latents across multiple viewpoints. We provide object-level evaluation metrics and a protocol for quantitatively assessing the accuracy of cross-view latent transport. We then demonstrate the utility of the resulting multi-view object representations in robotic manipulation and Gaussian Splatting, showing their applicability to both 3D scene representation and downstream tasks. Our code will be publicly available. Videos are available at https://sites.google.com/view/particletransport.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.