Where, When, and How Many Bits? Scaling Laws for SGD with Stochastic Rounding
Abstract
Low-precision training with stochastic rounding has become popular for large language models (LLMs). However, there is currently no theory to guide how much precision each matrix product needs or how it should be scheduled over training. We study one-pass SGD on random-feature regression with power-law spectra, where the features in the forward propagation and the weight-gradient products are stochastically rounded with time-varying precision. We derive an exact recursion for the rounded loss curve and scaling laws for the extra loss due to rounding. For the final weights of a training run, performed at a constant learning rate, these laws answer three questions. Where: rounding in the forward pass limits accuracy, while rounding in the weight-gradient product is hidden by the noise of SGD. When: the bits are best spent late in training. How much: the number of bits needed grows only logarithmically with compute budget, whereas a fixed precision lowers the exponent of the scaling law. Iterate averaging of the output and learning-rate decay remove the noise of SGD and thereby expose the rounding error in the weight-gradient product. This indicates that the learning-rate and precision schedules must be designed together.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.