acceptodds
Under review as a conference paper at ICLR 2027

SiamJEPA: On the Role of Siamese Student Encoders in JEPA

Abstract

Joint Embedding Predictive Architectures (JEPAs) have emerged as a promising framework for self-supervised representation learning by predicting latent embeddings of masked regions rather than reconstructing pixels. Existing JEPA methods typically employ a single student encoder, leaving the role of Siamese student encoders largely unexplored. In this paper, we propose Siamese JEPA (SiamJEPA), a JEPA framework with masked Siamese student encoders and an exponential moving average (EMA) teacher, which can also be viewed as a JEPA formulation of the brain-inspired representation learning model PhiNet. We further introduce Random Shuffle Teacher (RST), which removes spatial correspondence in teacher targets to encourage semantic patch representations, and develop an RST-based semantic-to-spatial curriculum. Experiments on ImageNet show that Siamese student encoders effectively regularize the JEPA objective, improving representation separability and accelerating early-stage learning. Moreover, under RST, stronger Siamese regularization substantially increases class-discriminative information in individual patch tokens, suggesting that the semantic bias arises from the interaction between RST and the Siamese objective rather than from shuffling alone. SiamJEPA consistently outperforms comparable single-encoder JEPA variants under limited training budgets. With RST-based curriculum learning and a ViT-Base backbone, SiamJEPA achieves 73.3% linear-probing accuracy after 450 epochs using a substantially simpler masking strategy, compared with 72.9% for I-JEPA after 600 epochs. Under I-JEPA's linear-probing protocol, the same checkpoint achieves 74.2%, outperforming I-JEPA (72.9%) and DSeq-JEPA (73.5%). These results demonstrate that Siamese student encoders provide an effective inductive bias for predictive representation learning and can be further enhanced through semantic-to-spatial curriculum learning.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.