DiJEPA: Self-Supervised Learning via Diffusion in Latent Space
Abstract
Joint-Embedding Predictive Architectures (JEPAs) learn semantic representations by predicting the latent embeddings of the masked target view from the visible context. Existing approaches use either direct regression in a continuous embedding space or token-softmax prediction over quantized targets. While effective, direct regression does not explicitly model a target distribution, whereas token-based prediction performs distribution matching in a discrete space; both rely on carefully designed distribution regularization to stabilize training and improve representation quality. We introduce DiJEPA, a JEPA-style framework that formulates visual representation learning as diffusion-based prediction over an online projection of the encoder's representation space. DiJEPA jointly trains a Feature Auto-Encoder (FAE) with the encoder, mapping high-dimensional representations into low-dimensional continuous latents whose distribution is optimized throughout representation learning. A diffusion predictor then reconstructs target-view latents from student-view representations through denoising, providing a distributional prediction objective in a continuous latent space. On ImageNet-1K with a ViT-L backbone, DiJEPA achieves 83.24% top-1 linear probing accuracy under a non-distilled setting, competitive with 83.3% from DINOv3. The jointly learned FAE latents can also be used for image generation, achieving 2.74 gFID without CFG in 80 epochs of latent diffusion training. Our results establish diffusion-based prediction over evolving FAE latents as a new predictive objective alongside direct regression and distribution matching over discrete tokens for self-supervised visual learning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.