acceptodds
Under review as a conference paper at ICLR 2027

On fast hyperparameter transfer under fixed and diverging time horizons

Abstract

Hyperparameter (HP) transfer, which avoid expensive tuning on large neural networks by reusing values tuned on small models, can substantially reduce large-scale pretraining costs. Yet existing evidence is largely qualitative and does not precisely characterize the computational gain. We study this methodology through the framework of fast transfer ghosh2025understanding, which relates computational benefits to the convergence rates of the optimal loss and HPs with model scale . We first prove a second-order Tensor Program master theorem for general neural computations at a fixed training horizon . For fully connected polynomial networks trained by gradient descent under the Maximal Update Parameterization (P), this yields a sharp fluctuation of the loss and optimal HPs. While this result quantifies loss and HP convergence, we argue that fixed- asymptotics do not explain the computational benefit of -transfer, as the leading centered loss fluctuations conflict with empirically observed scaling laws and loss monotonicity. This motivates us to investigate jointly diverging model width and training horizon in a minimal tractable setting: tuning the gradient descent learning rate in a power-law random features model. Using a resolvent expansion, we derive sharp loss asymptotics and HP convergence, and establish necessary and sufficient spectral conditions for fast transfer at width-dependent training horizons. We further discuss the scaling behaviors at the compute-optimal frontier and the effect of stochastic gradients.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.