acceptodds
Under review as a conference paper at ICLR 2027

RotateAttention-FP4: A Unified Training-Free and Data-Efficient QAT Framework for Quantized Attention

Abstract

Attention incurs substantial computation and memory traffic in video diffusion models, motivating 4-bit quantization for efficient inference. However, existing training-free 4-bit attention methods remain limited in both generation quality and runtime efficiency, while quantization-aware training (QAT) often incurs substantial training costs. We propose RotateAttention-FP4, a unified NVFP4 attention framework supporting accurate and efficient video diffusion through both training-free inference and lightweight adaptation. Our key insight is that activation outliers and stable channel-wise biases amplify quantization error, degrading the accuracy of training-free inference and increasing the burden on model adaptation. We refer to these distribution-induced errors as structural errors and suppress them through grouped Hadamard rotation and calibration-based channel centering using offline-estimated offsets. We further optimize quantization scales for both QKV and attention weights: rotation-aware QKV scale selection reduces FP4 rounding error, while probability amplification reduces underflow of attention-weight scales. Together, these corrections improve training-free inference accuracy and reduce the residual error that sensitive models must compensate for through lightweight QAT. Experiments on Wan2.2 and MiniMax H3 demonstrate generation quality close to 16-bit attention without any fine-tuning in conventional multi-step settings and on eight-step MiniMax H3. For the quantization-sensitive four-step Wan2.2 model, QAT with low-rank adaptation (LoRA) recovers generation quality using only 30 randomly sampled training examples. Our fused implementation further delivers a – end-to-end attention speedup over FlashAttention and – preprocessing throughput over SageAttention3.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.