acceptodds
Under review as a conference paper at ICLR 2027

Optimal Training-Time Scaling in Gradual Adaptation

Abstract

In gradual adaptation, how should the training time on each task change as the number of intermediate tasks increases? We study this question for overparameterized linear regression tasks that change smoothly and share a zero-loss solution. With tasks and training time on each, the final learning progress converges to a continuum curve when . The limiting progress is for small and for large , so both very short and very long training produce little progress. It follows that optimal per-task training times scale as , equivalently . Experiments on gradually rotated MNIST and a natural Yearbook time shift are consistent with less per-task training as the path is divided more finely.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.