Efficient Language Model Pretraining via Low-Rank Readout Modulation
Abstract
Efficient pretraining of language models (LMs) requires reaching a desired model quality with fewer optimization steps and minimal additional cost per step. Most work pursues this goal through data curation, improved optimizers, or more efficient Transformer variants. In this work, we focus on a less explored aspect: the LM's readout head. We introduce Low-Rank Readout Modulation (LRM), a simple, lightweight feature-wise gate applied once between the final normalization and tied LM head. LRM derives its scale from the embedding of the observed token, leaves the Transformer stack and KV cache unchanged, and requires neither another decoder pass nor a custom kernel. Across three seeds at 115M, 300M, and 1B scales, LRM improves held-out negative log-likelihood (NLL) on both standard Transformer and Multi-Head Attention Residuals (MHAR) backbones, with mean gains of 0.024–0.043 NLL. It adds at most 0.06% parameters to the base model and, at similar throughput, reaches loss targets with 15–30% fewer steps. Across six backbone–task pairs, parameter-efficient fine-tuning of models pretrained with LRM improves mean accuracy in five.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.