acceptodds
Under review as a conference paper at ICLR 2027

What Transfers from Non-Visual Procedural Warm-Up? Attention Operators as Structured Initializers for Vision Transformers

Abstract

Non-visual procedural warm-up can improve the subsequent training of Vision Transformers (ViTs), despite exposing the model to sequences with no visual or semantic overlap with images. What initialization structure does this warm-up leave behind? We study this question in a controlled setting: a ViT is warmed up for 15,000 steps on DYCK-64 and then trained for image classification. Through elimination interventions, we progressively discard warm-up parameters and isolate the bias-free query-key (QK) and output-value (OV) weight operators of selected attention blocks. We then refactor these operators into fresh projection matrices while randomly initializing the rest of the model, separating operator transfer from raw weight copying. Across three matched source/downstream seeds, late-block QK+OV initialization reaches 75.17 ± 0.30% on CIFAR-100 and 60.41 ± 0.64% on Tiny-ImageNet, compared with 74.74 ± 0.51% and 60.30 ± 0.70% for full warm-up transfer, retaining 121% and 103% of its gain over scratch. On ViT-Base/ImageNet-1K, expanding QK+OV coverage from the last four to the last 6/8/10/12 blocks yields 79.86/80.40/80.27/80.54% Top-1, versus 77.72% from scratch and 79.50% from full warm-up; with the source fixed, the last-eight setting reaches 80.32 ± 0.12% across three downstream seeds. Exact alternative factorizations further show that optimization can depend on factor coordinates even when the transferred operators are fixed. This behavior is source-dependent: under supervised ImageNet pretraining, the same QK+OV intervention recovers only 15% of the full transfer gain. Our results show that late-block attention-operator initialization can retain the benefit of a short procedural warm-up, while establishing a clear empirical boundary between procedural initialization and conventional visual feature transfer.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.