Freeze or Dither: Rounding-Aware Precision Floors for Low-Precision Training
Abstract
Convergence theory for low-precision training has produced only upper bounds, in a worst-case relative-error model blind to the rounding rule. We give the first lower bounds for training with floating-point weight storage, organized by one principle: a coordinate survives round-to-nearest (RN) when its update magnitude either exceeds the local grid spacing or carries noise at that scale, and freezes or floors when it does neither — freeze or dither. Sign-type updates can never dither. Below an explicit threshold they freeze bitwise, weight growth stops at a hard “binade wall,” and their movement-weighted noisy floor is exactly under both RN and stochastic rounding (SR), for arbitrary predictable step-size schedules, so at the stationary floor SR provably cannot help sign descent. SGD self-dithers: RN's rounding bias collapses as , which puts homogeneous-noise SGD–RN on the universal scale. But un-ditherable low-noise coordinates escape the rescue, and under heterogeneous noise with a shared constant step SGD–RN floors at , strictly worse than SGD–SR, so no rounding-blind analysis can be tight for both rules. A representation floor separately forces mantissa bits; with the sufficiency of tang2026convergence the requirement is . The floors are measured to their constants, and the wall is visible in real training: in a M-parameter GPT-2 storing weights in bf16 with no master copy, RN pins every LayerNorm gain at exactly its initialization of (a binade boundary) in of (optimizer, seed) cells, SR pins none; at the tuned budget the held-out RN/SR ratios order, in every seed, as ditherability predicts, and SR recovers the fp32 master's loss at a third of the weight memory. All code, logs and figures behind every number reported here are in the supplementary material and will be released publicly upon acceptance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.