LexDINO: Language-Grounded Prototypes for Visual Pretraining
Abstract
A general-purpose visual encoder should transfer effectively across recognition, dense prediction, and generative vision-language tasks. Existing pretraining approaches offer complementary strengths across these settings: visual self-supervision (e.g., DINO) is particularly effective for dense prediction, while image-text supervision (e.g., CLIP) often yields better transfer to generative vision-language models (VLMs). We introduce LexDINO, which brings language-derived semantic structure into DINO's visual self-distillation to improve vision-language transfer while retaining its superior visual capabilities. An independent caption corpus is clustered in an LLM's feature space, and the resulting centroids replace DINO's learned image-level prototypes. A second bank represents the same caption groups in a frozen image-text model's joint space, allowing its image tower to ground the corresponding DINO assignments. Semantic guidance therefore acts through DINO's own prediction space, alongside cross-view self-distillation and masked patch learning, without directly regressing backbone features or requiring captions paired with the training images. With the ImageNet-1k image corpus, backbone, and training schedule held fixed, LexDINO improves over a DINOv3 trained in the same way by 2.0 points in k-NN accuracy and 3.3 points in ADE20K mIoU. Under the same downstream VLM training recipe, it improves all twelve benchmarks, including a 25.4-point gain in the five-task OCR average, and its LLM-derived prototype bank outperforms a SigLIP2-text bank by 6.7 OCR points. At scale, a roughly 1B parameter LexDINO encoder, trained with a fraction of DINOv3's pretraining compute, achieves best or tied performance on all five dense benchmarks among the evaluated models below 2B parameters. Batch-independent sparse targets and full-grid patch supervision simplify its distillation into ViT-L/16 and ViT-B/16 students. Both students outperform size-matched DINOv3 encoders in recognition and VLM transfer while retaining strong dense prediction.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.