POST-HOC NORMALISATION: A PREREQUISITE FOR LINEAR EVALUATION OF SELF-SUPERVISED FOUNDATION MODELS
Abstract
Linear evaluation by fitting a linear classifier on frozen features is the standard yardstick for the quality of self-supervised representations. It is, however, acutely sensitive to a geometric artefact that has nothing to do with representational content: anisotropy, the concentration of features in a narrow, off-centre, correlated cone, which is severe in vision transformers and inflates apparent inter-class similarity. We show that the benefit of correcting this geometry post hoc, before the linear probe, is not a fixed property of the correction, but scales inversely with how well-conditioned the frozen features already are. We propose a supervised, training-free, representation-space normalization that isotropizes intra-class (noise) covariance while leaving inter-class (signal) structure intact, and position this normalisation precisely to distinguish it from in-network nor- malization (BatchNorm/LayerNorm), projector-space anti-collapse normalization (DINO centering), and unsupervised post-hoc whitening. On CIFAR-100, the correction lifts a reconstruction-pretrained backbone (BEiT) by +27.3 points, a self-distillation backbone trained on ImageNet-1k (≈1.3M images, DINO v1) by +12.2 points, an otherwise-identical backbone that adds a patch-level masked- image term at the same scale (iBOT) by +8.4 points, and a self-distillation back- bone trained on 142M curated images (DINOv2) by +1.2 points. This monotone ladder separates two contributing factors of anisotropy: at fixed data scale a stronger objective (DINO v1→iBOT), and at fixed objective term ≈ 110× more data (iBOT→DINOv2). Fairness of linear evaluation is clouded by anisotropy induced complexity of optimisation, which the proposed post-hoc normalisation mitigates. Anisotropy diagnostics on 30000 CIFAR-100 photographs and on ten other benchmarks corroborate the mechanism. We also demonstrate the merits of the normalisation on the task of organ classification in medical data analysis.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.