acceptodds
Under review as a conference paper at ICLR 2027

Foveation as a Self-Supervision Signal: Blurring Beats Blanking for Learning Robust Visual Representations

Abstract

The primate retina is strikingly non-uniform, with receptor density falling off precipitously from fovea to periphery. Every saccade delivers a different retinal sampling of the visual world, yet our visual systems relate multiple glimpses and assemble a coherent sense of the scene. At some level of representation, the visual system must therefore map different retinal views of the same scene onto similar neural codes. Here we hypothesize that learning to predict fine peripheral structure from foveal feedback at the next saccade provides a natural mechanism for self-supervised learning of good visual representations. We implement this by replacing blank masks of standard masked autoencoders (MAEs) with blur masks. A vision transformer (ViT) is pre-trained to reconstruct the full-resolution image from the corrupted input, and then fine-tuned on clean images for classification. We evaluate performance on CIFAR-10 and ImageNet. Blur-MAE beats blank-MAE on the pre-training reconstruction task. Critically, blur-MAE also outperforms blank-MAE on downstream classification, showing that it learns better generalizable representations. The gap of blur-MAE superiority widens under impoverished conditions such as higher mask ratios, fewer pre-training epochs, and smaller training sets. Band-limited noise probes reveal that blank-mask models rely on high-spatial-frequency (SF) under every impoverished condition while blur-mask models do not, explaining the classification gap. A closed-form linear-ridge encoder derived from the MAE objective approximates the empirical ViT SF tuning (Pearson ) and localizes the blank-vs-blur difference in the singular-vector structure of the corrupted-input matrix.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.