Fourier Positional Embeddings for Vision Transformers
Abstract
Convolutional neural networks (CNNs) have long shaped progress in computer vision by incorporating strong inductive biases such as translation equivariance and locality. Vision Transformers (ViTs) now reach comparable performance across many tasks, but rely on weaker built-in spatial priors and instead depend on positional encodings to represent image structure. We introduce Fourier Positional Embedding for Vision Transformers (FViT), a positional encoding inspired by Fourier analysis. Rather than adding positional embeddings to token representations, FViT represents relative spatial structure through learned complex-valued phase factors that modulate attention, preserving the flexibility of self-attention while better aligning positional encoding with image geometry. Empirically, FViT substantially outperforms absolute, relative, and sinusoidal positional embeddings and matches Rotary Positional Embeddings (RoPE) on CIFAR-100 and ImageNet. On COCO object detection, FViT substantially improves over the standard YOLOS baseline under a matched training setup and achieves higher average recall than RoPE across all reported recall metrics, while RoPE retains an advantage in average precision. Finally, we argue that FViT and RoPE perform strongly because their phase-based positional operations reflect how translations act through Fourier characters.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.