Weak Supervision Helps JEPA
Abstract
When a hand picks up a cup, understanding the action requires distinguishing the hand from the cup while capturing their interaction. Yet global video self-supervision leaves these object-level distinctions implicit. We show that weak object supervision during pretraining improves the quality and data efficiency of LeVJEPA representations. We use object masks generated offline by pretrained SAM2 to supervise an auxiliary instance-separation objective, bringing together patch features belonging to the same object while separating them from other objects and the background. Furthermore, we show that instance separation alone is the best-performing auxiliary objective, outperforming variants incorporating object-level SIGReg regularization and additional object-level alignment. The encoder architecture and retained-token budget remain unchanged. SAM2 is used only for preprocessing, and the auxiliary head is removed after pretraining, leaving no additional inference cost. Across matched experiments with ViT-Tiny, ViT-Small, and ViT-Large pretrained on 50K to 1M videos, our method improves over standard LeVJEPA at every evaluated scale. Identical attentive probes on frozen encoders achieve higher early-epoch and final accuracy. At 50K videos, Top-1 accuracy improves by approximately 9.4 percentage points on both Something-Something-v2 and Kinetics-400. The SSv2 advantage persists as pretraining data increases, reaching 6.7 points at 200K videos and 4.7 points at 1M. These gains also translate into improved data efficiency. On nested datasets, our model pretrained on 50K videos outperforms standard LeVJEPA pretrained on 95K by 6.1 SSv2 Top-1 points. In a complementary cross-scale comparison, our ViT-Small reaches 51.9% SSv2 accuracy on an SSv2-enriched 200K-video corpus, within 4.3 points of the ViT-Large baseline pretrained on 1M videos. Together, these findings identify local instance separation as an effective complement to global video self-supervision: a single auxiliary objective provides useful object structure during pretraining and enables stronger representations from limited data.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.