acceptodds
Under review as a conference paper at ICLR 2027

Optimizer Memory Schedules for Outscaling the Overtraining Axis

Abstract

We show that optimizer memory schedules can produce an increasing token-efficiency advantage over AdamW as Transformer training horizons grow. We compare AdamW, Muon, SOAP, and ADANA across four model sizes from 51M to 458M parameters and overtraining (OT) factors from 1x to 256x, sweeping the base learning rate at every setting. We introduce a timescale-based momentum cooldown rule that shortens optimizer memory during terminal learning rate decay, producing large long-horizon gains for ADANA that compound with log-time weight decay. With these schedules, ADANA outscales AdamW with a scaling advantage close to that predicted by DANA theory on power-law random features. Longer horizons generally favor longer fixed memory, but ADANA's scaling advantage persists after tuning AdamW's fixed memory separately at each horizon. Muon's token-efficiency advantage over AdamW is roughly constant across OT factors, while SOAP gains further at high OT. ADANA begins behind both matrix-preconditioned optimizers but surpasses Muon and becomes competitive with SOAP at our highest OT factors. These results establish training horizon as an essential axis for optimizer evaluation and design.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.