GaussRoPE: Rethinking 3D Rotary Encoding When Geometry Is Uncertain
Abstract
3D rotary position embeddings (RoPE) use camera poses and scene depths that can be uncertain. Position errors can distort the relative phases on which attention relies, while geometry refined inside the network cannot inform subsequent encodings. We introduce GaussRoPE, which connects rotary encoding, uncertainty propagation, and layerwise refinement through a unified structured Gaussian geometry state. Propagation maps pose and depth uncertainty into rotary phases, while cross-token shared camera errors are implemented through token-wise rotations. This shared-error sampling preserves the modeled phase correlations and remains compatible with standard attention, without requiring explicit pairwise uncertainty tensors. Visual features update the geometry and its uncertainty, enabling refined states to inform subsequent encodings. We evaluate GaussRoPE across five host models spanning novel-view synthesis, Gaussian reconstruction, 3D object detection, and depth estimation. Relative to the strongest competing encodings in our evaluation, the full GaussRoPE framework improves PETR mAP by approximately 28% under both original and perturbed inputs. Under geometric perturbations, it also improves MVSplat PSNR by 0.48 dB and reduces UniMatch AbsRel by 7.4% relative to competing methods under the same equally weighted pose/depth/joint protocol.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.