Effective Learning Rate Governs Loss Dynamics in Language Model Pretraining
Abstract
We establish ELR collapse as a quantitative empirical law in language model pretraining: learning rate (LR) and parameter norm govern loss dynamics primarily through their ratio, the effective learning rate (ELR). When tensor-wise ELRs are matched across runs, their loss trajectories collapse throughout training despite substantially different LRs and parameter norms. Across optimizers, architectures, datasets, and model scales, mean collapse errors are typically a few , below the seed-to-seed variation measured in a representative configuration. Systematic ablations identify normalization design and the timescale of LR–norm variation as key determinants of collapse precision. In standard training, tensor-wise ELRs approximately synchronize, following a shared temporal profile up to fixed amplitudes. Controlled interventions further show that adjusting only the base LR to match an aggregate ELR reproduces the loss dynamics induced by weight decay and Hyperball to high precision. Replacing LR with ELR enables a fitted functional scaling law (FSL) to transfer across norm-control methods without refitting. The resulting FSL also explains delayed acceleration, where norm control initially raises loss but improves it later, and guides the design of an ELR schedule that further lowers final loss. Together, these findings identify ELR as an intrinsic clock rate governing loss dynamics across LR schedules and norm-control methods.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.