Efficient Dense Visual Representation Learning via Multi-Granularity Supervision
Abstract
Visual representations encode the semantic and visual content of real-world images in a feature space. They are widely used in semantic segmentation, cross-modal retrieval, and Multimodal Large Language Models (MLLMs). However, existing vision encoders, such as CLIP, are predominantly trained via image–text contrastive learning. This approach relies heavily on large-scale web-crawled alt-text as supervision, resulting in low data-efficiency during training. Moreover, the global instance-level contrastive objective limits the model's ability to capture fine-grained semantics and their corresponding relationships. In this paper, we propose TRIDENT, a data-efficient vision encoder pretrained with complementary forms of supervision: (i) instance-level image–text contrastive learning with a fine-tuned LLM as the text encoder to achieve cross-modal alignment and transfer its rich world knowledge to the vision encoder; (ii) region-level feature distillation to capture semantic consistency and spatial structure; and (iii) pixel-level reconstruction to preserve local visual details. Together, these objectives provide complementary global, regional, and pixel-level supervision, enabling richer dense visual representations and more data-efficient pretraining. Trained on only 1B samples, TRIDENT broadly outperforms eight competitive vision encoders trained on 4–40B samples across image–text retrieval, dense prediction, and downstream MLLM benchmarks. Compared with SigLIP2, TRIDENT improves I2T/T2I retrieval by 12.2/18.2 R@1 on Urban1K and 17.8/20.9 R@1 on ShareGPT4V, while achieving gains of 4.8 mIoU on ADE20K and 6.1 mIoU on PASCAL VOC for semantic segmentation. On MLLM downstream tasks, TRIDENT outperforms SigLIP2 on 10 of 11 benchmarks while matching it on MMMU, including gains of 3.4 points on VQAv2, 3.0 points on GQA, and 7.3 points on MMBench-EN. We will release the corresponding code and model checkpoints to facilitate further research in visual representation learning.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.