acceptodds
Under review as a conference paper at ICLR 2027

When Schedule-Free Fails: Online Learning Rates for Anytime Training

Abstract

Schedule-free (SF) optimization avoids a predefined training horizon, yet its performance can worsen with continued training at a fixed learning rate. We introduce a new view of SF based on the pending update, which accumulates new updates before they are gradually applied to the output iterate. This view motivates **ORA, Online Learning Rates for Anytime Training**, which selects learning rates to control the size of new updates relative to the pending update. Our local analysis in an aligned subspace of a deterministic two-layer linear model shows that SF-Muon generally converges to a nonstationary point while ORA-Muon converges to an optimum. In noisy two-layer regression, smaller fixed learning rates delay SF-Muon's risk increase, while ORA closely tracks the lower envelope of SF-Muon curves. In 124M language model pretraining, ORA reaches the same validation loss as SF-Muon with **42-51% fewer tokens**. Compared with SF-Muon, ORA more closely follows the empirical loss scaling law under cosine decay and yields more accurate extrapolation. These results support studying anytime methods through scaling laws rather than only at a single horizon. ORA also matches the loss of strong Muon cosine baselines with **23% and 33% fewer tokens** on 124M and 720M models, respectively. Extrapolation suggests that these token savings grow further with longer training.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.