acceptodds
Under review as a conference paper at ICLR 2027

Growing to See: Efficient Pretraining of Self-Supervised Vision Transformers

Abstract

Progressive depth growth builds large networks by starting with shallow models and inserting layers during training. Recent advances in large language models have renewed interest in this longstanding approach, demonstrating benefits for training efficiency and reasoning alongside changes in internal computations. However, how growing affects representation quality and computational patterns in self-supervised vision transformers remains unclear. In this work, we systematically investigate progressive depth growth in two popular self-supervised families, namely masked autoencoders (MAE) and DINOv3, and analyze how growing affects downstream performance, intermediate representations, and internal computations. Across global and dense tasks, we show that growing improves downstream performance while reducing training FLOPs by 13.9–24.8%, with ImageNet linear-probe accuracy gains of up to 3.36 percentage points. Growing also improves downstream performance at intermediate layers and reshapes internal computational patterns, introducing recurring alternations between global and localized attention in MAE and earlier local-to-global attention progression in self-distilled DINOv3. Together, these findings establish progressive depth growth as an effective and efficient pretraining strategy for improving downstream performance across global and dense visual tasks with less computation, and advance our understanding of how growing shapes visual computation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.