Reconstruction is not Perception: A Simple Proof of Necessity but Insufficiency
Abstract
Masked autoencoders (MAE) provide a scalable approach to visual representation learning, yet strong classification performance often requires fine-tuning, richer probes, or longer pretraining. We ask whether this mismatch reflects downstream underdetermination: models with the same reconstruction loss can nevertheless have substantially different downstream utility. We study this question through reconstruction-loss level sets in the joint encoder–decoder parameter space. Controlled sweeps of ViT-B/16 MAEs on ImageNet-100 reveal substantial variation in linear-probe accuracy among checkpoints at matched reconstruction loss, with the variation generally increasing as reconstruction loss decreases. We derive a first-order local characterization of this phenomenon and adapt a classification-guided predictor–corrector traversal that improves or degrades downstream performance while keeping reconstruction loss nearly fixed. Across nine decoder–loss settings, these walks span – percentage points of accuracy, – the matched-loss pretraining spread, and extend beyond the pretraining accuracy range in every setting. The phenomenon also transfers to public ImageNet-1k checkpoints, where walks improve ViT-B/16 and ViT-L/16 by and percentage points at nearly unchanged reconstruction loss. Mechanistically, reconstruction and classification gradients remain nearly orthogonal throughout pretraining, while walk-induced spectral changes concentrate in a late-learned tail containing only – of target variance. These results show that reconstruction loss does not determine downstream representation quality and identify directions that reconstruction weakly constrains yet matter substantially for classification.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.