acceptodds
Under review as a conference paper at ICLR 2027

What Ages a Gradient? When Optimizer Time and Gradient-Field Relevance Diverge

Abstract

AdamW assigns age by counting optimizer updates, but equal update counts can correspond to very different amounts of learning-state change. We ask whether elapsed optimizer time is sufficient to predict how well a historical gradient direction remains aligned with the current gradient field. Fixed held-out probes keep text identity constant, isolating gradient-field evolution from minibatch variation. On Pythia-160M, elapsed steps and accumulated learning rate leave a reproducible predictive gap: OBJECTIVE, a loss-normalized gradient–update work diagnostic, reduces held-out prediction error by 9.87% and 7.48% in two environments and transfers between them without refitting. The gap survives a new seed and both shorter and longer first-moment memory; under the longer-memory setting, a scale-normalized parameter-motion diagnostic reaches +16.00%. A post-discovery decomposition shows that learning state captures substantially more of the residual: a four-variable state summary raises held-out from 0.588/0.621 to 0.741/0.731. Parameter motion and training progress explain part of this signal, while temporal misalignment erodes the remainder. For actual stochastic gradients, relevance also reflects alignment at gradient birth: unconditional state transfer is weak, while a prespecified birth-conditioned comparison is positive. Together, these results show that fixed-timescale optimizer memory compresses training histories that can have different present-day gradient relevance because they traverse different learning states.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.