acceptodds
Under review as a conference paper at ICLR 2027

How Far Can MXFP4 Go in Pretraining?

Abstract

While pretraining large language models in 4-bit floating point (FP4) can substantially reduce training cost, FP4 formats struggle to cover the dynamic range of the weights, activations, and gradients encountered during training. Popular FP4 formats therefore pair each small block of values with a shared scale. MXFP4 uses power-of-two block scales that are simpler to support in hardware, while NVFP4 spends mantissa bits on its local scale to reduce quantization error. The two formats thus trade hardware simplicity against scale precision, and recent work reports larger pretraining degradation for MXFP4 than for NVFP4. In this work, we propose a MXFP4 pretraining recipe, MXFP4-MARS, based on recent advanced techniques, including Hadamard rotation, stochastic rounding, and Overflow-Aware Scaling (OAS). Our recipe also leverages Macro Block Scaling (MBS), which maintains higher-precision scale at a coarser granularity while keeping hardware-friendly power-of-two per-block scales. At matched perplexity, our MXFP4 recipe closes the token-overhead gap to NVFP4 from the previously reported 36% down to 0.21% on average, a 171 reduction in additional tokens required. Also, at the same training budget, a 7B model trained using MXFP4-MARS yields a 1.4% higher average downstream score than the model trained with NVFP4. These results identify scale hierarchy as a key design dimension for low-precision training and show that MXFP4 can be quality-competitive with NVFP4 while keeping the hardware-friendly power-of-two scales, making it an attractive, portable target for FP4 pretraining.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.