SigmaTransfer: Uncertainty Transfer from Small to Large Networks under muP
Abstract
Reliable predictive uncertainty in Laplace approximations depends critically on the prior precision, yet selecting it requires a posterior sweep that is prohibitively expensive for neural networks with billions of parameters. Under the Maximal Update Parametrization, we derive a rescaling of the prior covariance that stabilizes its induced prior kernel as model width grows. This leads to SigmaTransfer: we select the precision on a smaller model and zero-shot transfer it to the much larger model, i.e., without searching for the precision on the larger model at all. Under explicit conditions, we show convergence of the prior kernel, posterior covariance, selected precision, and posterior-derived decisions, and verify SigmaTransfer across regression, image classification, and Transformer readouts. For example, measured precision-sweep speedups reach when transferring from width 128 to 4096 on MNIST, at a target-NLL degradation of ; transferring from a public 1B to 7B model gives , with a degradation of NLL. The same posterior stability also enables transfer of acquisition, OOD-detection, and abstention decisions without constructing a target posterior.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.