acceptodds
Under review as a conference paper at ICLR 2027

Beyond the Central Ray: Ray Differential Attention for Novel View Synthesis

Abstract

Generalizable novel-view synthesis reconstructs unseen views from posed images using transformer in a single forward pass. Transformers tackle this task by aggregating features across views, where performance critically depends on how attention encodes camera geometry. Existing methods predominantly represent each token as an infinitesimal central ray, discarding the patch's local angular scale, anisotropy, and shear across the viewing sphere. To address this limitation, we introduce Ray Differential Attention (RDA). First, we represent each token by a ray differential—a unit ray along with its spatial image derivatives capturing the local patch footprint without requiring scene depth. Second, RDA incorporates this geometry into both attention routing and feature aggregation. For routing, dual-primal query–key transforms yield relative, world-frame-invariant attention scores. For aggregation, an invertible value transform enables consistent cross-view content transport via an exact query-side inverse. This establishes a broader principle that camera geometry should actively govern both token interaction and feature expression, rather than serving as passive conditioning. In a controlled LVSM host, RDA outperforms five camera baselines across PSNR, SSIM, and LPIPS on RealEstate10K and DL3DV. Rigorous ablations confirm that both the full ray differential footprint and the dual read/write pathways are essential to these gains.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.