S-JEPA: Self-Sparsifying Representations in Joint-Embedding Predictive Architectures
Abstract
Rectified LpJEPA enables direct control over representation sparsity in joint-embedding predictive architectures (JEPA), yet its performance degrades as representations approach extreme sparsity. We show that much of this degradation is an optimization cost rather than a limit of the sparse code, as aggressive sparsity shrinks gradients and limits early formation of informative features. We therefore introduce S-JEPA, a self-sparsifying JEPA that paces sparsification from loss curvature and corrects optimization strength from code density, adapting both throughout learning without labels or reference runs. At sparsity, reference-free S-JEPA improves frozen encoder and sparse projector accuracy over fixed-sparse training by up to and points, and exceeds its non-sparse counterpart on CIFAR-100 and STL-10, while retaining of non-sparse accuracy on ImageNet-100. Its gains over fixed-sparse training hold across ResNet and ViT backbones and four optimizers. Pretrained on CIFAR-100, S-JEPA improves mean transfer accuracy across five datasets by points over fixed-sparse training, matching non-sparse transfer, and outperforms dense JEPA codes in nearest-neighbor retrieval with less storage.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.