Pretraining Induces Stable Spectral Subspaces for Downstream Task Adaptation
Abstract
Finetuning pretrained models often exhibits a low-dimensional structure in parameter space, with prior work focusing on where adaptation occurs. We instead investigate which parameter subspaces remain stable across finetuning and downstream tasks. We analyze pretrained weight matrices across vision and language models and find that leading spectral subspaces remain largely unchanged during finetuning, with stability consistently stronger in earlier layers. This stability emerges progressively during pretraining and increases with pretraining dataset size, suggesting that downstream adaptation largely preserves the leading spectral structure formed during pretraining. We further show that downstream adaptation can be restricted to pretrained leading spectral subspaces in deeper layers, achieving performance close to full finetuning while updating only 0.2% of model parameters on GLUE. Our findings offer a new parameter-space perspective on how pretraining shapes downstream adaptation, highlighting leading spectral subspaces as stable structure inherited across tasks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.