acceptodds
Under review as a conference paper at ICLR 2027

On the Critical Batch Size for Low-Precision Training

Abstract

Large language models are increasingly trained with large batches and low-precision arithmetic to improve computational efficiency. The critical batch size (CBS) marks the onset of diminishing returns from larger batches, yet how this threshold changes with numerical precision remains poorly understood. We study the CBS of low-precision training through theoretical scaling laws for quantized stochastic gradient descent in sketched linear regression. Our analysis replaces the vector- and matrix-level error-alignment constraints of prior multiplicative quantization models with more practical bounds on the relative variance of independent, unbiased coordinate-wise rounding. Our theory predicts that lower precision increases the critical batch scale while preserving its scaling exponent with the data budget. The analysis explains this shift through a balance between quantization-induced noise and the cost of fewer optimization updates, accounting for parameter-rounding noise shared across a batch. We combine controlled regression experiments with matched-quality language-model comparisons to test the predicted scaling and its qualitative transfer to pretraining. This analysis motivates joint selection of numerical precision and batch size to exploit the useful range of data parallelism.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.