acceptodds
Under review as a conference paper at ICLR 2027

MASaLA: Mass-Aware Fusion of Sparse and Low-rank Attention

Abstract

The quadratic computational cost of dense attention remains a central obstacle to training Transformers on long inputs. Efficient approximations typically restrict each query to a subset of keys or compress global interactions through low-rank and kernel-based representations. Sparse methods preserve sharp, localized interactions but may miss diffuse global structure, whereas low-rank methods capture broad mixing but can blur concentrated interactions. We propose MASaLA, a mass-aware fusion of sparse and low-rank attention operators that retains dense query, key, value, and output projections. A block-sparse branch computes exact attention over selected key blocks, while a low-rank branch summarizes diffuse global interactions. To account for the branches' different normalization masses, MASaLA uses mass-aware fusion that rescales the sparse contribution according to its estimated relative attention mass. The framework computes attention at subquadratic cost without materializing the full attention matrix. We further show that including each query's own block in the sparse support gives the hybrid operator the capacity to attain full rank regardless of the low-rank budget, overcoming the rank restriction of a purely low-rank approximation. We evaluate MASaLA on long-context classification benchmarks spanning 16K–64K tokens.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.