A Theoretical and Experimental Analysis of the Robustness of Similarity-based Positional Encoding Under Rotation
Abstract
Transformer architectures heavily rely on positional encodings (PEs) to inject spatial structure; however, standard approaches remain highly sensitive to spatial rotations. In this paper, we present a unified theoretical and empirical investigation into the rotational robustness of similarity-based positional encoding (simPE). Theoretically, we prove that while simPE is not strictly rotation-invariant, it exhibits provable stability under simple Lipschitz regularity assumptions, yielding explicit Frobenius norm perturbation bounds that cleanly decouple architectural features from rotational magnitude. Empirically, we evaluate models trained on canonical orientations against increasing rotation angles across five diverse datasets, by comparing simPE with a comprehensive suite of baseline PEs. Across the tested datasets, simPE provides better performance (measured in terms of classification accuracy and F1-score) in small-to-moderate rotation regimes (), with respect to competing PE methods. Crucially, on a real clinical benchmark (Chest X-Ray), simPE outperforms all competing PEs across every evaluated angle (), and allows the corresponding ViT architecture to reach significant levels of accuracy. Our findings bridge theoretical stability guarantees with practical spatial robustness, establishing simPE as a highly effective encoding strategy for domains subject to routine geometric perturbations.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.