Where the FP4 Pretraining Gap Sits, and Why That Decides What Transfers Across Scale
Abstract
Training a language model in 4-bit floating point (FP4) is faster than in BF16 but ends at a higher loss. Recipes narrow this gap by keeping some computations, or the end of training, in higher precision, and these choices are usually tuned on small models. We ask where the gap comes from in one FP4 recipe at three model sizes, and find that the answer changes with scale. In a small model nearly all of the gap comes from rounding in the forward pass; as the model grows this share falls, and rounding the gradients, alone and together with the forward pass, takes over. At every size a similar share of the gap is the cost of the rounding still in progress, which a switch to BF16 removes at once, while the rest stays in the weights. Since higher precision in one part of training removes at most the share of the gap that arises there, where the gap sits decides what transfers across scale: a higher-precision forward pass removes most of the gap in a small model but less than half in a large one, while a short switch to BF16 at the end removes the same share at every size. A few runs at the target scale measure these shares before any recipe is trained.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.