Self-Supervised Representation Learning as Conditional Generation
Abstract
Self-supervised representation learning (SSL) typically learns from view consistency or masked modeling by matching, reconstructing, or predicting specified targets. However, the relationship between an observation and its target can inherently be one-to-many: multiple semantically valid representations or visual contents may be plausible given the same observation, yet existing SSL objectives do not explicitly model this conditional uncertainty. We introduce CondGen, a new formulation of SSL as conditional generation in representation space, which models the conditional distribution \(p_\theta(z_i\mid z_j)\) between representations of different views using diffusion models. By explicitly modeling this distribution, CondGen captures the one-to-many relationship across views rather than enforcing consistency with a specified target. To further incorporate global semantic structure, we introduce Decoupled Sinkhorn Clustering, which learns balanced semantic prototypes independently of encoder optimization and uses them to refine target representations for conditional generation. Extensive experiments show that CondGen learns robust, semantically structured, and spatially coherent representations across diverse tasks, including shape/structural robustness, image generation, video object segmentation, and 3D correspondence estimation. In particular, CondGen even significantly surpasses DINOv3 on the unsupervised object discovery task, despite its substantially larger-scale pretraining, demonstrating the potential of conditional representation modeling as a new paradigm for scalable self-supervised learning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.