acceptodds
Under review as a conference paper at ICLR 2027

Optimizer Memory: Persistence and Intervention Value in Language Model Adaptation

Abstract

Optimizer moments preserve gradient history, but a persistent causal effect need not identify a useful intervention. Controlled language-model adaptation separates intervention at a reached checkpoint from planning at a common warm origin. Across twelve allocations, late second-moment replacement reduces target loss by 674.54 micro-NLL while increasing reference loss by 15.60; the complete procedure nevertheless loses to higher-rate retained-state training. A separate comparison finds a small refresh benefit of 34.02 micro-NLL, but additional ordinary updates reduce target loss below an old-rich intervention when its training-time allowance is reinvested. On an unedited native trajectory, recent-history reweighting worsens both domains in all eight confirmation allocations after symmetric finite-menu suffix selection. Native graft comparisons clarify this failure: both the magnitude change and the direction change raise target loss, although their reference effects differ. Actual adaptation to WikiText-103 retains the recent-direction target penalty in one further setting. These paired interventions distinguish causal persistence, conditional marginal gain and useful continuation without establishing a general preference for older statistics or practical optimizer superiority.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.