acceptodds
Under review as a conference paper at ICLR 2027

Self-Supervised Representation Learning as Conditional Generation

Abstract

Self-supervised representation learning (SSL) typically learns from view consistency or masked modeling by matching, reconstructing, or predicting specified targets. However, the relationship between an observation and its target can inherently be one-to-many: multiple semantically valid representations or visual contents may be plausible given the same observation, yet existing SSL objectives do not explicitly model this conditional uncertainty. We introduce CondGen, a new formulation of SSL as conditional generation in representation space, which models the conditional distribution \(p_\theta(z_i\mid z_j)\) between representations of different views using diffusion models. By explicitly modeling this distribution, CondGen captures the one-to-many relationship across views rather than enforcing consistency with a specified target. To further incorporate global semantic structure, we introduce Decoupled Sinkhorn Clustering, which learns balanced semantic prototypes independently of encoder optimization and uses them to refine target representations for conditional generation. Extensive experiments show that CondGen learns robust, semantically structured, and spatially coherent representations across diverse tasks, including shape/structural robustness, image generation, video object segmentation, and 3D correspondence estimation. In particular, CondGen even significantly surpasses DINOv3 on the unsupervised object discovery task, despite its substantially larger-scale pretraining, demonstrating the potential of conditional representation modeling as a new paradigm for scalable self-supervised learning.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.