A Matter of Direction: Eigenframe Distribution Matching for Self-Supervised Learning
Abstract
Self-supervised objectives seek representations that preserve informative variation while avoiding collapse. Yet a representation can distinguish images while losing variation along some directions, a partial form of collapse. Why might such an objective respond weakly to this loss? We investigate this question in methods that match feature distributions through random one-dimensional projections. Each projection mixes many directions, allowing variation in healthy directions to mask a deficiency in another. In a local Gaussian model, we show that this effect becomes stronger as the representation dimension grows. Averaging more projections reduces sampling noise but does not strengthen the expected response. We call this effect spectral dilution. This observation motivates VARTE, a self-supervised objective that replaces random directions with the covariance eigenvectors of features averaged across augmented views. These directions separate the representation's modes of variation, allowing the matching objective to examine each one directly. We analyze this choice theoretically and test its consequences through controlled interventions and image representation learning. On ImageNet-1K, VARTE reaches 74.7% linear-probe accuracy with a ViT-B/16 and transfers competitively across recognition and specialized-domain tasks. Our work highlights a simple principle: preventing collapse depends not only on the distribution we ask features to match, but also on the directions along which we measure them.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.