What Transfers from Non-Visual Procedural Warm-Up? Attention Operators as Structured Initializers for Vision Transformers
Abstract
Non-visual procedural warm-up can improve the subsequent training of Vision Transformers (ViTs), despite exposing the model to sequences with no visual or semantic overlap with images. What initialization structure does this warm-up leave behind? We study this question in a controlled setting: a ViT is warmed up for 15,000 steps on DYCK-64 and then trained for image classification. Through elimination interventions, we progressively discard warm-up parameters and isolate the bias-free query-key (QK) and output-value (OV) weight operators of selected attention blocks. We then refactor these operators into fresh projection matrices while randomly initializing the rest of the model, separating operator transfer from raw weight copying. Across three matched source/downstream seeds, late-block QK+OV initialization reaches 75.17 ± 0.30% on CIFAR-100 and 60.41 ± 0.64% on Tiny-ImageNet, compared with 74.74 ± 0.51% and 60.30 ± 0.70% for full warm-up transfer, retaining 121% and 103% of its gain over scratch. On ViT-Base/ImageNet-1K, expanding QK+OV coverage from the last four to the last 6/8/10/12 blocks yields 79.86/80.40/80.27/80.54% Top-1, versus 77.72% from scratch and 79.50% from full warm-up; with the source fixed, the last-eight setting reaches 80.32 ± 0.12% across three downstream seeds. Exact alternative factorizations further show that optimization can depend on factor coordinates even when the transferred operators are fixed. This behavior is source-dependent: under supervised ImageNet pretraining, the same QK+OV intervention recovers only 15% of the full transfer gain. Our results show that late-block attention-operator initialization can retain the benefit of a short procedural warm-up, while establishing a clear empirical boundary between procedural initialization and conventional visual feature transfer.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.