From Pixels, Without Pre-training: Joint Generative and Self-Supervised Representation Learning in One Model
Abstract
Strong image generation models today are conditioned on class labels, aligned to frozen pretrained encoders, or built on separately trained autoencoders or tokenizers. While effective, the cost is that generation becomes contingent on supervision and pretraining: labels must be annotated, and encoders, autoencoders, and tokenizers must be pretrained for the target domain. This paper studies **joint generative and self-supervised representation learning in one model**, such that generation is self-conditioned without any labels or pretrained models. This is challenging since the two objectives are mismatched: contrastive learning consumes clean augmented views and keeps coarse, invariant semantics, while flow matching consumes noisy images and must preserve the fine detail and spatial layout that contrastive learning discards. We propose a framework Self-conditioned Generation on Self-supervised representation (SCION), whose core is a single pixel-space encoder conditioned on both the diffusion timestep and a conditioning embedding. The encoder serves two roles that differ in what that embedding is. For representation learning, the conditioning embedding is one learned global vector shared by all images, and the encoder’s [CLS] token yields the semantic embedding trained by the contrastive loss. For generative training, the conditioning embedding is the [CLS] embedding of the image itself, and the encoder’s patch tokens pass through a decoder to predict the image. To sample without a reference image at inference, we also learn a prior over the embedding in the same run. With gradient-norm balancing and stop- gradient tricks, all losses are jointly optimized in a single run. The resulting method is fully self-supervised, with no pretrained model and no labels. On ImageNet 256 × 256, with the same JiT-B/16 recipe and without representation guidance, SCION reaches 8.92 FID, surpassing the class-unconditional baselines iREPA, which aligns to a pretrained DINOv2 (46.44), and RCG, which conditions on it (14.27). Under the JiT-L/16 recipe, SCION achieves the best pixel-space class-unconditional FID of 5.89 and 3.47 without and with representation guidance, outperforming RCG with the ADM recipe (6.24).
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.