MXAttention: Data-Free Optimal Scaling and Pre-Normalization Quantization for MXFP4 Attention
Abstract
The quadratic cost of attention limits diffusion-based video generation. MXFP4 attention offers a promising path toward lower-cost inference, but direct quantization often degrades generation quality. We identify two numerical causes: power-of-two shared scaling creates a clipping–underflow trade-off within MXFP4 blocks, while quantization in the online softmax loop breaks row-wise normalization, so the attention weights no longer sum to one. We propose MXAttention, a data-free post-training quantization framework with two components. First, Universal Optimal Scaling (UOS) exploits the periodic structure of power-of-two microscaling to minimize a global MXFP4 quantization-error objective. It yields the closed-form, distribution-independent boundary Qmax = 7.25, without calibration or per-layer search. Second, Pre-Normalization Quantization (PNQ) uses the same quantized exponential tiles in the normalizer and output updates, preserving row-wise normalization. On Wan2.2 and HunyuanVideo, MXAttention substantially improves frame-level similarity to FP16 outputs. It preserves FP16-level generation quality with less than 0.01 absolute degradation on all reported VBench metrics and achieves video quality competitive with strong NVFP4-based baselines. We further evaluate LLM quantization on Qwen3-14B, observing accuracy gains for attention quantization, weight–activation quantization, and their combination. On one Ascend 950-series NPU, MXAttention delivers these accuracy gains at nearly the same attention latency as the PNQ-enabled MXFP4 baseline (0.13% difference). MXAttention retains a 2.70× attention speedup and a 1.84× total video-generation task speedup over BF16 attention.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.