Compute Less, Capture More: Improving Sparse Attention with Mixed Precision
Abstract
Sparse attention reduces computation by selecting which query–key interactions to evaluate, but typically computes every selected interaction at the same numerical precision. We find that this is inefficient: importance remains highly uneven within the selected set, while useful interactions are still omitted. We introduce mixed-precision sparse attention, which converts this excess precision into broader coverage. It preserves high precision for important interactions, lowers the precision of less important retained interactions, and uses the saved compute to recover interactions that the original sparse method would otherwise skip. By reusing the geometry or importance ranking of existing sparse methods, it expands coverage without a new selection mechanism or additional compute. Our analysis shows that this exchange can reduce attention-output error because broader coverage better preserves dense attention distributions and its benefit can outweigh the error introduced by lower precision. We further develop a fused kernel that executes BF16, FP8, and FP4 interactions within a single attention operation with little overhead. Across video generation and long-context language understanding, mixed-precision sparse attention consistently improves the quality–compute frontier of existing sparse attention methods.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.