acceptodds
Under review as a conference paper at ICLR 2027

Grounding Self-Supervised Latent Novel View Synthesis with Cross-View Consistency

Abstract

A self-supervised novel view synthesis objective offers a promising approach to learning latent scene representations from unlabeled images. Despite remarkable progress, inferring geometrically grounded representations from unposed multi-view images remains a fundamental challenge. Latent representation methods achieve strong photorealism but suffer from shortcut learning, and fail to recover the underlying camera motion. Explicit-representation methods resolve this ambiguity by rendering through 3D Gaussians, but their non-convex optimisation converges only under a carefully scheduled training curriculum. In this work, we ask whether geometry can be encoded into features while preserving the expressive capacity of a latent space. We introduce a feed-forward model that bakes geometry into features through a photometric reprojection objective. Our proposed framework comprises a permutation-equivariant camera estimator, a ray-conditioned renderer, and a dense depth decoder that facilitate reconstructing every view through a warping operation via predicted depth, intrinsics, and relative pose. We evaluate both photorealism and raw pose accuracy under interpolation and extrapolation settings. Experiments show that, in the extrapolation setting, our method outperforms prior self-supervised methods in PSNR, LPIPS, and strict raw pose accuracy.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.