acceptodds
Under review as a conference paper at ICLR 2027

LeVJEPA: Efficient & Scalable Video Pretraining without the Heuristics

Abstract

Video carries the temporal structure of the physical world, yet learning representations from it remains computationally expensive. Current self-supervised methods either prevent representation collapse through architectural asymmetries (an exponential-moving-average target encoder, a stop-gradient, and a predictor), or avoid it by reconstructing masked content in pixel space with a dedicated decoder. We introduce LeVJEPA, the first extension of LeJEPA to video. A single encoder is trained with an invariance loss between global and local views of a clip, while collapse is prevented by SIGReg with provable guarantees. Thus, neither architectural asymmetries nor pixel reconstruction is needed. This formulation has two properties that we study in this paper. First, the pretraining cost depends only on the number of tokens the encoder observes. Dropping tokens uniformly at random reduces this number and, at the same time, improves downstream accuracy. At matched epochs on identical data, LeVJEPA matches or outperforms V-JEPA 2 across ViT-S/B/L with to less pretraining compute. At matched FLOPs, it exceeds the strongest video baseline by points on ImageNet-1K while remaining competitive on motion-centric benchmarks. Second, since no target encoder or predictor is involved, the encoder can be trained with block-causal attention without loss in accuracy, so that the representation of each frame depends only on past frames. Finally, compared to DINOv2 trained on frames of the same videos with matched compute, LeVJEPA remains within points on ImageNet-1K while nearly doubling the accuracy of DINOv2 on Something-Something-v2. These results suggest that, once its computational overhead is removed, video is a viable substrate for general-purpose visual pretraining.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.