SCOPE: From Analytical Scale Oracles to Candidate Scale Search for NVFP4 Quantized Training
Abstract
Recent hardware support for microscaled FP4 formats has made NVFP4 a compelling target for reducing the cost of large language model pre-training. However, existing NVFP4 training recipes still show a noticeable quality gap to BF16, primarily due to forward weight-activation quantization under the sparse FP4 codebook and shared block scale. We propose **SCOPE**, a forward-centered NVFP4 quantization framework that combines an analytical scale oracle with a deployable candidate search. SCOPE first introduces an analytical per-block scale oracle based on FP4 rounding breakpoints to characterize the reconstruction improvement recoverable by adaptive scale selection. For deployment, SCOPE replaces online breakpoint enumeration with a compact fixed-candidate search, preserving data-dependent scale selection while keeping the online path simple. Backward gradient-to-quantization-noise diagnostics further show that lightweight SR/Had-SR estimators remain well above the theoretically motivated diagnostic reference in our setting, allowing us to use a low-overhead backward path and focus the main optimization effort on forward scale selection. On NanoChat pre-training up to 1.9B parameters and 38B tokens, SCOPE reduces the BF16 BPB gap to 1.01%, a relative gap reduction of \(~30%\) over Quartet-II and \(~47%\) over the NVIDIA NVFP4 recipe. In a matched 500M experiment, SCOPE reduces the measured complete-recipe training time by 16.5% relative to Quartet-II while achieving better validation quality. Additional PTQ and pre-training experiments show transfer across dense LLMs, MoE models, and vision-language settings.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.