acceptodds
Under review as a conference paper at ICLR 2027

ThriftAttention: Selective Mixed Precision for Long-Context FP4 Attention

Abstract

Efficient attention algorithms are critical to mitigate the quadratic cost of attention in long-context workloads. Such methods often rely on low-bit quantisation and/or hard sparsity to accelerate inference. FP4 attention on Blackwell GPUs offers large speedups but degrades sharply relative to FP16 as context length grows while sparsity methods discard query-key interactions outright which can degrade model quality. We address both failure modes through ThriftAttention, a mixed-precision attention mechanism that uses a heuristic to promote only the most important interactions to FP16 while computing the remainder in FP4. ThriftAttention relies on our finding that the output impact of quantisation error is highly non-uniform and increases with the importance of each query-key interaction, concentrating functionally relevant error in a small number of attention blocks that contain the most important tokens. Our experiments show across long-context benchmarks and model families that by computing only 5% of the attention blocks in FP16, ThriftAttention recovers on average 92.3% of the FP4FP16 performance gap. We demonstrate ThriftAttention's advantage as context length grows, the regime where FP4 attention performs worst.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.