Generation-induced Representation Learning: Self-Guided Diffusion without Pretext Tasks
Abstract
Generative training alone can drive semantic representation learning. We demonstrate this via a new framework that jointly trains an image encoder and a conditional diffusion model end-to-end from scratch, using no labels, no contrastive loss, no masking, and no data augmentation beyond horizontal flips. The encoder outputs a latent vector that both controls conditional generation and serves as a suitable feature representation for downstream semantic tasks. An additional diffusion model trained over the encoder's output space allows for sampling of latent representations, and, chained together with the conditional model, yields a system for unconditional image generation. Key to our system's learning dynamics are independently sampled noise timesteps for the conditional model's image and latent inputs. On CIFAR-10, CIFAR-100, and ImageNet-100, our system substantially improves unconditional generation over an architecturally matched DDPM baseline with identical compute. Our central claim is one of existence and minimality: an encoder trained purely by the denoising diffusion loss, under a deliberately minimal flip-only augmentation budget, reaches kNN classification accuracy comparable to dedicated self-supervised methods such as SODA, despite using no pretext task whatsoever. Moreover, equipping our method with the same data augmentation recipe as SimCLR, so that the two differ only in their training objective, our purely generative objective outperforms SimCLR's contrastive objective on three of four probing protocols across CIFAR-10 and CIFAR-100. Together, these results establish that conditional generation, when coupled with appropriate training mechanisms, is itself a viable representation learning objective.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.