Raster-Free View Synthesis with Point-Anchored Geometric Embeddings
Abstract
Geometry foundation models (GFMs) such as VGGT estimate camera poses and depth from multi-view images in a single forward pass, and recent feed-forward novel-view synthesis (NVS) methods build on this geometry by rasterizing it into the target view. However, the estimated geometry inevitably contains small errors, and rasterization turns them directly into visible artifacts, such as objects appearing as overlapping copies. We propose RaViS, a raster-free view synthesis model built on a simple principle of attending to the geometry rather than rasterizing it. RaViS synthesizes the target view with a transformer and uses the estimated geometry only to tell attention how image patches are related in 3D. To this end, we introduce a point-anchored geometric embedding that places each patch at its own estimated 3D point, so that attention between two patches depends on their relative 3D position and camera rotation. Since the geometry only steers attention and never enters the output image, its errors merely shift where the model attends instead of appearing as artifacts. To further inject 3D structure and semantics, we feed the synthesis transformer with the latent features that the GFM already computes while estimating the geometry. Using only GFM-estimated camera poses, RaViS outperforms prior feed-forward methods on RealEstate10K and DL3DV and in zero-shot generalization. On RealEstate10K, it even outperforms similarly sized baselines that have access to ground-truth camera poses.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.