Scalable Patch-Level Self-Supervised Learning
Abstract
Self-supervised learning (SSL) at scale produces powerful visual representations. However, most scalable SSL methods rely on ad hoc combinations of multiple objectives and stabilization mechanisms. Taking a step back, we ask if we can design a high-performing, yet principled SSL algorithm. Starting from the multi-view assumption–stipulating that task-relevant content is captured by the information common to different views–we construct an information-theoretic objective decomposing into interpretable terms. This derivation yields JEM, a student–teacher method that learns by aligning corresponding patch representations across views, explicitly regularized by information and structure preservation losses. JEM trains stably from 300M to 7B parameters, and, to our knowledge, is the first latent-space patch-level method demonstrated at 7B scale. Across all scales, JEM reaches strong performance on both global and dense probing tasks, on segmentation benchmarks consistently surpassing the DINOv2 algorithm–an influential foundation for today’s strongest visual SSL methods. Notably, at 7B parameters, it also exceeds DINOv3's performance on panoptic segmentation, despite DINOv3 being trained on a 10x larger dataset with several refinement stages. These results demonstrate that we can indeed build an SSL algorithm which is principled and stable, without sacrificing downstream performance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.