What Have Vision Models Learned in the Deep Learning Era?
Abstract
What have vision models learned in the deep learning era, and how general are their representations? We systematically revisit the decade of visual self-supervised learning (VSSL) and geometry, generation, and video models, with references including supervised models, vision-language models (VLMs), multimodal world models (MMWMs), and frontier models. Through over 2,500 training and evaluation experiments, we distinguish original-setting comparisons, controlled pre-training, and common downstream evaluation on the Basic Five: image classification, object detection, segmentation, depth estimation, and video classification, complemented by an extended transfer suite. All scores are computed from our own environment, never copied from prior reports. Our main findings connect the past, present, and future directions of visual representation learning. (Past) Over the past decade, vision-native representations have become substantially more general, supporting transfer across semantics, dense prediction, and geometry, although progress and model strengths remain task-dependent. (Present) Modern vision-native representations, exemplified by DINOv3, outperform the evaluated VLM and MMWM references on multiple visual tasks, demonstrating broad transfer from vision-native learning itself. (Future direction) These achievements are not fixed endpoints: learning-design changes improve both classical and modern VSSL approaches, while scaling yields further gains across all five core tasks. Moreover, a separate matched-subset comparison with GPT-6 Astra reveals complementary strengths: DINOv3 performs better on ADE20K semantic-point accuracy and comparably on NYUv2 relative-depth accuracy. Overall, these results connect historical progress to the current frontier and future directions, providing an empirical account of what vision models have learned and a foundation for developing more general vision-native representations.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.