Predicting the Predictable: Feature Suppression in Joint-Embedding Predictive Architectures
Abstract
Self-supervised methods that train an encoder to map different augmented views of an image to similar embeddings must prevent the trivial solution in which all embeddings become identical. Recent work does so by constraining the embedding distribution itself: LeJEPA requires embeddings to follow an isotropic Gaussian and proves that this distribution minimizes worst-case downstream risk. However, we observe that fixing the distribution turns the prediction objective into a competition for a bounded variance budget. We find that this competition orders directions by the fraction of their variance that is stable across views—an intraclass correlation—rather than by how easily a feature is learned. Whichever feature ranks higher occupies the budget and suppresses the other. Because the constraint fixes each direction’s variance rather than which feature fills it, numerical rank does not drop even as effective rank, which measures how evenly variance is used, does. In contrast, I-JEPA, which predicts masked representations rather than comparing views, is unaffected by a competing feature that substantially degrades LeJEPA. To address this, we propose a penalty that removes variance stable across views of an image but not shared with other images, leaving the isotropy constraint intact. Empirically, we show that it reduces feature suppression and improves downstream performance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.