acceptodds
Under review as a conference paper at ICLR 2027

Mixed Precision Scaling Laws for Quantized Pretraining

Abstract

Since scaling laws were introduced for modern LLMs kaplan2020scaling, several works have extended them to account for quantization, tracking the emerging low-precision data types supported on AI accelerators. We develop a more accurate scaling law for quantized pretraining, grounded in optimization theory and validated empirically. Our formulation separates the effects of quantization on model capacity and on optimization efficiency. Through convergence analysis, we further observe that both forward and backward quantization can introduce an optimization error that persists with continued training, which we model explicitly in our law. To design our experimental regimes used to fit and evaluate our scaling law, we first revisit existing 4-bit training recipes and decompose them into their constituent forward and backward pass quantization choices, isolating the components that most strongly affect loss. Our Qwen3 experiments under selected 4-bit regimes, together with published Llama-style experiments, show that our law predicts loss with up to 11× lower MSE than existing laws as model size and token count scale. Finally, we extend our law to layer-wise mixed precision, and identify the optimal quantized-pretraining recipe, i.e., model size, training tokens, and per-layer forward/backward precision, under a hardware-aware compute budget.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.