Learnable Multipliers: Freeing the Scale of Language Model Matrix Layers
Abstract
Language-model pretraining typically combines weight decay (WD) on parameter matrices with learnable scales, such as RMSNorm gains. Scales that can be absorbed into adjacent matrices add no representational capacity, yet learning them separately often improves training. We untangle the roles of WD and separately learned scales through a single mechanism: optimizer updates can drive unchecked norm growth, WD restrains this growth but holds matrix norms near potentially suboptimal values, and separately learned scales let the effective matrix scales adapt to the data. We demonstrate this mechanism both empirically, through controlled experiments with Adam and Muon, and theoretically, using the Adam-trained LM head as a case study in which second-moment normalization suppresses the signal from rare informative gradients. We introduce learnable multipliers (LRMs) to extend this scale adaptation systematically across model matrices. LRMs learn diverse feature scales across and within layers while remaining compatible with P learning-rate transfer. Progressively adding scale degrees of freedom across model matrices reduces training loss. In long pretraining across architectures and optimizers, our full LRM configuration delivers substantial downstream gains comparable to the Adam-to-Muon switch.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.