acceptodds
Under review as a conference paper at ICLR 2027

A Lasting Bias of Initialization: How Initialization Shapes Generalization in Vision Transformers

Abstract

Initialization is typically designed to stabilize optimization, yet its effects can persist throughout training. Recent work for example showed that warming up Vision Transformers (ViTs) on procedural text, such as Dyck languages, surprisingly improves downstream image classification despite seeing no images during warm-up. We investigate which changes to initialization led to this. We find that much of the improvement can be recovered by matching a small set of layerwise parameter and activation statistics, including effective weight scales, LayerNorm statistics, low attention entropy and negative MLP pre-activations in early-to-middle layers, and large residual write ratios in late layers. We extract a total of 108 of these scalar statistics across the layers of procedurally warmed-up reference models and show that imposing them on a randomly initialized ViT is sufficient to match or exceed the improvements from procedural warm-up. Models trained from procedural and our modified initializations also show persistently lower intermediate-layer prediction accuracy and memorize fewer corrupted labels. These observations are consistent with an implicit regularizing effect. Together, our results show that a few initialization statistics can reproduce much of the benefit of procedural warm-up and provide evidence that initialization can impose a lasting inductive bias on learning and generalization.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.