NADA: Unlocking Fast and Accurate NVFP4 PTQ through Distortion-Aware Optimization
Abstract
NVFP4 provides a wide effective dynamic range while retaining hardware-native FP4 computation on NVIDIA Blackwell GPUs, enabling high quantization fidelity with efficient low-precision inference. However, accurate NVFP4 post-training quantization (PTQ) remains challenging, particularly for activations. Simply suppressing outliers does not guarantee lower activation quantization error, largely because the nonuniform E2M1 codebook induces uneven distortion. We identify that NVFP4 activation reconstruction is governed by three interdependent and sequentially coupled decisions at the tensor, channel, and block levels. This paper proposes an NVFP4 Activation Distortion-Aware (NADA) optimization framework for fast and accurate NVFP4 PTQ. NADA consists of tensor-level Distortion-Aware Activation Channel Clustering (DACC), channel-level Bounded Iterative Top- Rescaling (BITR), and block-level Analytical Block-Scale Optimization (ABSO). DACC and BITR are calibrated offline to reorganize activation channels and selectively rescale sensitive channels, respectively, while ABSO performs a search-free analytical scale update for dynamic activation blocks at runtime. Across eight models ranging from 135M to 70B, NADA achieves state-of-the-art accuracy while preserving high end-to-end inference efficiency. Code is available at https://anonymous.4open.science/r/NADA-5B92.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.