LeVLJEPA: End-to-End Vision-Language Pretraining Without Negatives
Abstract
Vision-language encoders are increasingly used as frozen backbones for multimodal language models and dense prediction, which consume the grid of patch tokens rather than a single pooled embedding. Image-text pretraining objectives, however, are defined on the pooled embedding, so the patch tokens are shaped only indirectly. We introduce LeVLJEPA, the first fully non-contrastive end-to-end vision-language pretraining method. Image and text embeddings predict each other through modality-specific predictors with stop-gradient targets, and collapse is prevented by regularizing each modality with SIGReg. Thus, no negatives, temperature, momentum encoder, or teacher-student schedule are needed. On Datacomp-L, LeVLJEPA matches CLIP and SigLIP under linear probing and is more robust to background shifts, while its zero-shot accuracy is lower. On the patch tokens, which no objective supervises directly, it improves over the strongest contrastive baseline by more than mIoU on both semantic segmentation benchmarks. Used as the frozen visual backbone of a vision-language model, it obtains the highest accuracy on GQA, VQAv2, and POPE with two different language models. These results suggest that non-contrastive image-text pretraining is a viable way to learn visual features for dense prediction and multimodal language models.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.