HROPE: HOMOGRAPHY RELATIVE POSITIONAL ENCODING
Abstract
Geometric Relative Positional Encoding has become the primary mechanism for injecting spatial inductive biases into multi-view transformers. Existing designs condition attention on camera-level frustums or ray-marched depth, associating cross-view tokens through the sensor geometry. Such sensor-centric formulations, however, describe only where a token is observed from, leaving the rich surface geometry of the scene that the token depicts unmodeled. In this paper, we advocate a shift toward surface-aware positional encoding that explicitly grounds patch tokens on the physical surfaces. We introduce HRoPE, which represents the local geometry of each patch as a slanted 3D plane, following the slanted support-window assumption of classical multi-view stereo. We inject token-level, plane-induced homographies into attention as relative transformations. This explicitly enables the transformer to rectify perspective distortions and align patch features across viewpoints rather than having to learn this alignment implicitly from data. By construction, this design preserves the global frame invariance of prior relative encodings, while explicitly enabling the transformer to rectify perspective distortions and align patch features across viewpoints rather than having to learn this alignment implicitly from data. Our encoding requires no ground-truth geometry: the plane parameters can be predicted end-to-end under a view-synthesis objective alone, or derived from foundation models. Extensive experiments on RealEstate10K, CO3D, and DL3DV show that HRoPE consistently outperforms the evaluated geometric positional encodings for novel view synthesis. Extensive ablations and analysis shows the effectiveness of the proposed encoding.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.