acceptodds
Under review as a conference paper at ICLR 2027

GLARE: GENERALIZABLE LENS-AGNOSTIC RELATIVE-RAY ENCODING FOR 3D RECONSTRUCTION

Abstract

Feed-forward 3D reconstruction models inherit strong geometric priors from perspective imagery, but these priors are tied to how image coordinates parameterize viewing directions. When the camera projection changes, the same image-grid relation no longer corresponds to the same observation geometry, limiting the transfer of pretrained multi-view reasoning across lens configurations. We introduce GLARE (Generalizable Lens-Agnostic Relative-ray Encoding), which formulates projection adaptation as conditioning latent multi-view interactions on image-inferred camera-local viewing geometry. For each image, we estimate a continuous field of viewing directions and derive local ray descriptors that characterize how visual tokens sample the camera’s projection. Rather than treating these rays as scene points or imposing them as output constraints, we use their relative descriptor phases to modulate feature comparison and aggregation across views. This separates observation geometry, which describes how each image samples the scene, from scene geometry and camera motion, which remain jointly inferred from visual evidence. The resulting formulation enables a perspective-pretrained reconstruction prior to reason over diverse central-camera projections without requiring calibration, camera poses, or camera-family labels at inference. We therefore shift lens adaptation from camera-specific output modeling to projection-aware latent reasoning, providing a unified mechanism for reusing pretrained multi-view geometric priors across heterogeneous camera systems.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.