acceptodds
Under review as a conference paper at ICLR 2027

FreeAttn: Unlocking Native E3M4 for Efficient and Accurate Attention

Abstract

The quadratic cost of attention limits long-context prefill and video generation. Low-bit quantization increases throughput but introduces numerical error. At the same eight-bit width, E3M4 trades a narrower dynamic range for finer normal-number spacing than E4M3. We independently identify and validate native packed F2FP conversion and matrix multiply–accumulate (MMA) modes for E3M4 on Blackwell SM120, and connect them into the native attention path of FreeAttn. The framework relates QK/PV quantization errors to attention outputs under shared scaling, organizes composable corrections through attention identities, and reduces their cost through shared statistics and fused execution. Matched-format controls show lower aggregate BF16-relative probability and pixel error with E3 than E4. Evaluated for both fidelity and latency, optimized all-FP8 E3 accelerates Qwen3-8B/32B prefill by 1.15–1.16× and Wan2.1-14B video delivery by 1.43× over Torch SDPA. The FP4 path with Basic correction achieves 1.45–1.51× complete-attention-call speedup over Torch SDPA on seven common shapes, with lower aggregate operator output error than same-bitwidth SageAttention3. These results establish native E3M4 as a useful representation and execution option for attention inference.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.