acceptodds
Under review as a conference paper at ICLR 2027

Self-Supervised Encoders Learn a Task Prior in Their Spectrum That Isotropy Erases

Abstract

Self-supervised objectives increasingly push representations toward an isotropic Gaussian, which LeJEPA proves optimal for worst-case linear probes. We show that once tasks come from a distribution, a representation's spectrum is a label-free prior over them, and isotropy is the uniform prior. A one-penalty ridge probe on spectrum is generalised ridge on the latents with per-direction penalty , and is the Bayes (Wiener) estimator iff , the prior's variances, for any sample size, dimension and design that excites every direction (Lindley–Smith read from the representation's side), and a power-law prior recovers the spectral-bias learning-curve rate. Theorem A prices a spectrum in labels, and Theorem B shows that a -class task's prior has rank , so isotropy flattens exactly the directions that carry the task. Under a protocol on 23 encoders and 15 vision and text tasks (169 cells), isotropy costs a median 21 accuracy points at 200 labels (range 0.5–64), and recovering the discarded prior takes a median 500 labels, a price predicted () by the cost of detecting the weakest task direction, found post hoc on 20 cells, then replicated on the other 149. The prior comes back without labels. Averaged over other tasks, it recovers a median 89% of the oracle gain, and training toward a non-flat spectral target instead of isotropy raises 200-label accuracy from 60% to 80%, with the isotropy-trained encoder still worse even given the oracle prior. A supervised control pays the same price, so the effect is a property of linear probing itself.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.