Fourier Output for Visual Coordinate Regression
Abstract
Coordinate regression is a fundamental component of many computer vision tasks, yet most existing methods rely on direct scalar regression, coordinate discretization and classification, or task-specific output representations. We investigate whether a common representation can improve coordinate prediction across diverse visual tasks. While Fourier encodings are widely used to represent coordinate inputs, their use as representations for coordinate outputs remains underexplored. We introduce Fourier Output, which represents each scalar coordinate using multi-frequency sine–cosine features and directly supervises predictions in the Fourier encoding space, while leaving the underlying backbone unchanged. To efficiently recover coordinates at inference time, we develop a hierarchical decoding algorithm that progressively resolves phase ambiguity from low to high frequencies without candidate search. Through a controlled patch localization study, we systematically analyze the representation, supervision, frequency configuration, and decoding strategy, establishing a practical design for Fourier Output. We then evaluate it on human pose estimation, homography estimation, optical flow, and vision-language model (VLM)-based localization. Across these diverse tasks, Fourier Output consistently improves prediction accuracy over direct coordinate regression with minimal architectural changes, while training curves show that it can reach comparable accuracy with fewer training updates. These results demonstrate that Fourier features, widely used for representing coordinate inputs, can also provide a simple and general representation for visual coordinate outputs.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.