acceptodds
Under review as a conference paper at ICLR 2027

Feedforward Novel View Synthesis using Latent Representations for In-the-Wild Images

Abstract

Recent feed-forward models have shown promise for novel view synthesis (NVS) from unconstrained photo collections. However, existing training setups offer limited coverage of real-world appearance variation, transient occlusion, or diverse source-view configurations, while differences in evaluation protocols complicate comparisons. We introduce the Wild-NVS dataset and benchmark, curated from internet photo collections, with a sampling strategy that constructs source–target view sets using bidirectional covisibility, joint target coverage, and viewpoint diversity. The benchmark provides scene-level splits and fixed evaluation protocols across multiple source-view counts. We also propose Wild-NVS, a feed-forward model that synthesizes target views from latent representations. It combines bottlenecked appearance conditioning with a learned transient prediction head that guides source-feature suppression before decoding. Experiments on Wild-NVS and PhotoTourism demonstrate state-of-the-art performance among deterministic feed-forward methods for in-the-wild NVS, with improved appearance modeling and reduced transient artifacts. Our models and dataset will be made public.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.