When Does Learning-Rate Retuning Help? Attainable Output Geometry Under a Fixed Update Budget
Abstract
Learning-rate schedules are often tuned for one training setting and reused when models are adapted to new data under fixed update budgets. A finite-step optimization trajectory can itself act as a statistical regularizer. We study when retuning the schedule after a task change improves on preserving the nominal finite-step filter. We characterize the leading retuned-minus-preserved risk through the statistical distance between the oracle drift and the one-sided attainable output set. For distinct interior steps and a simple spectrum in dimensions, the attainable correction space has dimension . Some spectral corrections are therefore inaccessible when , whereas the unrestricted commuting optimum becomes locally attainable when . At equal steps, step splitting generates first-order output corrections missed by the parameter Jacobian, where indexes the local task change. For two steps, this singular geometry is the endpoint of a near-equal transition governed by , where is the nominal half-gap. These results explain why unrestricted spectral retuning can overestimate attainable corrections while local Jacobian sensitivity can underestimate them. Synthetic experiments support the theoretical predictions, while real-data diagnostics distinguish preservation, schedule reuse, and retuning.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.