acceptodds
Under review as a conference paper at ICLR 2027

Exp-Reduced Two-Level FP8 Attention for Vision Transformers

Abstract

As GPU compute throughput increases, softmax becomes an increasingly significant bottleneck in attention on modern GPUs, particularly those based on NVIDIA's Blackwell architecture. Existing approaches either sacrifice accuracy, as often observed with sigmoid- and ReLU-based softmax substitutes, or fail to deliver wall-clock speedups in practice. This work proposes the Exp-Reduced Two-Level (ERTL) function, a precision-agnostic alternative to softmax for vision transformers (ViTs). The ERTL function uses block statistics and shifted-ReLU weights for intra-block normalization while retaining exponential normalization across blocks. This design substantially reduces the number of exponential evaluations and maps naturally to tiled attention kernels with an appropriate block size. Furthermore, an FP8 attention kernel is developed by fusing the ERTL function with lazy output rescaling. Extensive experiments on ViTs show that ViTs using the ERTL function recover near-baseline accuracy with only lightweight fine-tuning. Moreover, the FP8 attention kernel using the ERTL function achieves up to a 1.55× speedup over BF16 exact-softmax attention and increases throughput by up to 245.5 TFLOPS compared with FP8 exact-softmax attention.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.